EVEMISSLAB
English

結果RST-2026-0101v0.1

三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×

各條件平均(可用試驗的品質/wall 秒/裝置焦耳):A0 0.5/10.21/1715.6;A1 0.5/9.87/1819.0;A2 0.5/79.09/15143.0;A3 0.6667/99.13/15921.5;A4 0.6667/69.08/13546.2;A5 0.5/31.91/5321.5。能量相對 A0:A1 1.06、A2 8.827、A3 9.281、A4 7.896、A5 3.102。36 次試驗共量得 GPU 能量 0.0891 kWh、試驗時間 29.9 分鐘;記憶體駐留從 A0 的 72.1 GiB·s 升到 A3 的 714.6 GiB·s。

研究狀態
STABLE 目前的研究結論相對穩定
證據等級
E2 受控實驗
結果
MIXED
資料基礎
REAL MODEL 真的跑了語言模型;模型、版本與設定都記在頁面上。
版本
0.1
更新
2026-09-07
建立
2026-09-03
領域
Evaluation
計畫
PRG-2026-0101 智能的物理計量(IPM)
作者
Neo.K (EveMissLab)
AI 協作
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT) — 2026-09-07 diagnostic, Splice (Claude Code, Anthropic) — execution and RESULT note

觀察到的結果

metrics
ssr
1.0
sdr
0.0
scm
device_energy_j
3.1019
wall_time_s
3.1264
by_condition
A0
quality_mean
0.5
quality_n
4
success_rate
0.5
wall_time_s
10.21
device_energy_j
1715.6
energy_ratio_vs_A0
1.0
gpu_peak_memory_gib
7.23
gpu_memory_residency_gib_s
72.1
gpu_utilization_integral_s
6.57
A1
quality_mean
0.5
quality_n
4
success_rate
0.5
wall_time_s
9.87
device_energy_j
1819.0
energy_ratio_vs_A0
1.06
gpu_peak_memory_gib
7.32
gpu_memory_residency_gib_s
70.0
gpu_utilization_integral_s
6.26
A2
quality_mean
0.5
quality_n
4
success_rate
0.5
wall_time_s
79.09
device_energy_j
15143.0
energy_ratio_vs_A0
8.827
gpu_peak_memory_gib
7.28
gpu_memory_residency_gib_s
570.0
gpu_utilization_integral_s
54.09
A3
quality_mean
0.6667
quality_n
3
success_rate
0.6667
wall_time_s
99.13
device_energy_j
15921.5
energy_ratio_vs_A0
9.281
gpu_peak_memory_gib
7.34
gpu_memory_residency_gib_s
714.6
gpu_utilization_integral_s
70.02
A4
quality_mean
0.6667
quality_n
3
success_rate
0.6667
wall_time_s
69.08
device_energy_j
13546.2
energy_ratio_vs_A0
7.896
gpu_peak_memory_gib
7.28
gpu_memory_residency_gib_s
496.9
gpu_utilization_integral_s
46.93
A5
quality_mean
0.5
quality_n
4
success_rate
0.5
wall_time_s
31.91
device_energy_j
5321.5
energy_ratio_vs_A0
3.102
gpu_peak_memory_gib
7.24
gpu_memory_residency_gib_s
228.7
gpu_utilization_integral_s
22.17
marginal_yield
A0->A1
allocated_device_time_s
device_energy_j
0.0
wall_time_s
A1->A2
allocated_device_time_s
device_energy_j
0.0
wall_time_s
0.0
A2->A3
allocated_device_time_s
device_energy_j
0.00021406648843853404
wall_time_s
0.008315919866432606
A3->A4
allocated_device_time_s
device_energy_j
wall_time_s
A4->A5
allocated_device_time_s
device_energy_j
wall_time_s

詮釋

interpretation
The cost side of the scaffolding response curve is real and steep; the quality side is flat because the tasks were already solved at A0 and because the quality axis was confounded (RST-2026-0102). This is one point on the 'no gap' side of F4 with almost no weight: it neither supports nor refutes the scaffolding-separation hypothesis on non-trivial tasks. The A3/A4 0.667 is not a gain: the aborted CON-003 trials dropped out of the quality denominator and the surviving mean rose.

限制

limitations
  • quality_n is 4 (A0–A2, A5) or 3 (A3, A4) per condition because CODE-001 quality is unavailable and three trials aborted — means over 3–4 values.
  • Device-measured GPU energy at ~24 % sampling overhead; not marginal energy.

主張

來源關係目標狀態ID
RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×qualifiesTHY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯ACTIVEREL-2026-0486
RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×qualifiesCLM-2026-0104 F4——鷹架分離ACTIVEREL-2026-0487

關係

來源關係目標狀態ID
RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×qualifiesTHY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯ACTIVEREL-2026-0486
RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×qualifiesCLM-2026-0104 F4——鷹架分離ACTIVEREL-2026-0487
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03)producesRST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×ACTIVEREL-2026-0485
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象qualifiesRST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×ACTIVEREL-2026-0490

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/results/RST-2026-0101/
機器可讀
/ai/results/RST-2026-0101/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports