EVEMISSLAB
English

結果RST-2026-0102v0.1

診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象

嚴格輸出契約合規 12/36(MATH-003 12/12、CODE-001 0/12、CON-003 0/12),但非 canonical 的事後檢查發現所有完成的選中輸出都正確——數學 12/12 canonical correct、程式 11/11 selected outputs pass all hidden tests after outer Markdown fence removal、約束 10/10 selected outputs satisfy all constraints after format-only normalization——整體 33/33。啟用驗證器的試驗 18 次,驗證器協定失敗 3 次(A3 2/6、A4 1/6、A5 0/6),每次都在已生成八個候選之後。工具呼叫 0 次、重試 0 次;24 個在中止時被拋棄的候選在 schema 裡既非選中也非丟棄。品質可得率在聚合檔報 22/36、在 protocol_compliance.csv 報 33/36。遙測取樣占 wall time 的 0.244。

研究狀態
STABLE 目前的研究結論相對穩定
證據等級
E2 受控實驗
結果
NEGATIVE
資料基礎
REAL MODEL 真的跑了語言模型;模型、版本與設定都記在頁面上。
版本
0.1
更新
2026-09-07
建立
2026-09-03
領域
Evaluation
計畫
PRG-2026-0101 智能的物理計量(IPM)
作者
Neo.K (EveMissLab)
AI 協作
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT) — 2026-09-07 diagnostic, Splice (Claude Code, Anthropic) — execution and RESULT note

觀察到的結果

metrics
strict_protocol_compliance
count
12
total
36
rate
0.3333333333333333
posthoc_semantic_diagnostic
canonical
false
math
12/12 canonical correct
code
11/11 selected outputs pass all hidden tests after outer Markdown fence removal
constraint
10/10 selected outputs satisfy all constraints after format-only normalization
completed_selected_outputs_correct
33/33
verifier_failure
count
3
verifier_enabled_trials
18
rate
0.16666666666666666
by_condition
A3
2/6
A4
1/6
A5
0/6
operational_totals
model_invocations
186
trajectories
162
retries
0
tool_calls
0
verifier_passes
18
candidates_created
162
candidates_selected
33
candidates_discarded
105
failure_events
6
candidates_abandoned_on_abort
24
quality_availability_reporting
aggregate_analysis
22/36
protocol_compliance_csv
33/36
inconsistency
true
instrumentation
mean_sampling_call_wall_fraction
0.2436
target_sampling_ms
250
observed_cadence_ms
335–350
instrument_revisions_required_before_XA-07
  • R1 typed subject/system failure is valid data, separate from instrument failure
  • R2 split task_semantic_quality / output_contract_compliance / system_completion_reliability / verifier_protocol_reliability
  • R3 failure-aware aggregation (no survivor means)
  • R4 abandoned-candidate accounting
  • R5 preserve verifier parse diagnostics
  • R6 cheaper telemetry (persistent nvidia-smi / NVML)
  • R7 tool-trigger tasks

詮釋

interpretation
Negative for the instrument, informative for the theory. Task semantic quality ≠ protocol/serialization compliance — the model solved everything it completed and was scored 0.5 on average for not obeying a JSON/no-fence contract; same-model verification reduced system reliability rather than raising quality; tool access ≠ tool utilization (enabled, never used). The bundle stays immutable as XA-06 empirical v0.1; the next step is to revise the instrument and re-run a small validation set, not to re-roll failures until they disappear.

限制

limitations
  • The post-hoc semantic check is explicitly non-canonical and does not make the strict outputs compliant.
  • One model; whether stronger models obey the output contract is unknown.

主張

來源關係目標狀態ID
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象supportsTHY-2026-0106 結構化品質、hard gate 與規格—驗證分離ACTIVEREL-2026-0489
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象qualifiesRST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×ACTIVEREL-2026-0490
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象qualifiesTHY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯ACTIVEREL-2026-0491

關係

來源關係目標狀態ID
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象supportsTHY-2026-0106 結構化品質、hard gate 與規格—驗證分離ACTIVEREL-2026-0489
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象qualifiesRST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3×ACTIVEREL-2026-0490
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象qualifiesTHY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯ACTIVEREL-2026-0491
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03)producesRST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象ACTIVEREL-2026-0488

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/results/RST-2026-0102/
機器可讀
/ai/results/RST-2026-0102/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports