結果RST-2026-0102v0.1
診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象
嚴格輸出契約合規 12/36(MATH-003 12/12、CODE-001 0/12、CON-003 0/12),但非 canonical 的事後檢查發現所有完成的選中輸出都正確——數學 12/12 canonical correct、程式 11/11 selected outputs pass all hidden tests after outer Markdown fence removal、約束 10/10 selected outputs satisfy all constraints after format-only normalization——整體 33/33。啟用驗證器的試驗 18 次,驗證器協定失敗 3 次(A3 2/6、A4 1/6、A5 0/6),每次都在已生成八個候選之後。工具呼叫 0 次、重試 0 次;24 個在中止時被拋棄的候選在 schema 裡既非選中也非丟棄。品質可得率在聚合檔報 22/36、在 protocol_compliance.csv 報 33/36。遙測取樣占 wall time 的 0.244。
觀察到的結果
metricsstrict_protocol_compliancecount- 12
total- 36
rate- 0.3333333333333333
posthoc_semantic_diagnosticcanonical- false
math- 12/12 canonical correct
code- 11/11 selected outputs pass all hidden tests after outer Markdown fence removal
constraint- 10/10 selected outputs satisfy all constraints after format-only normalization
completed_selected_outputs_correct- 33/33
verifier_failurecount- 3
verifier_enabled_trials- 18
rate- 0.16666666666666666
by_conditionA3- 2/6
A4- 1/6
A5- 0/6
operational_totalsmodel_invocations- 186
trajectories- 162
retries- 0
tool_calls- 0
verifier_passes- 18
candidates_created- 162
candidates_selected- 33
candidates_discarded- 105
failure_events- 6
candidates_abandoned_on_abort- 24
quality_availability_reportingaggregate_analysis- 22/36
protocol_compliance_csv- 33/36
inconsistency- true
instrumentationmean_sampling_call_wall_fraction- 0.2436
target_sampling_ms- 250
observed_cadence_ms- 335–350
instrument_revisions_required_before_XA-07- R1 typed subject/system failure is valid data, separate from instrument failure
- R2 split task_semantic_quality / output_contract_compliance / system_completion_reliability / verifier_protocol_reliability
- R3 failure-aware aggregation (no survivor means)
- R4 abandoned-candidate accounting
- R5 preserve verifier parse diagnostics
- R6 cheaper telemetry (persistent nvidia-smi / NVML)
- R7 tool-trigger tasks
詮釋
interpretation- Negative for the instrument, informative for the theory. Task semantic quality ≠ protocol/serialization compliance — the model solved everything it completed and was scored 0.5 on average for not obeying a JSON/no-fence contract; same-model verification reduced system reliability rather than raising quality; tool access ≠ tool utilization (enabled, never used). The bundle stays immutable as XA-06 empirical v0.1; the next step is to revise the instrument and re-run a small validation set, not to re-roll failures until they disappear.
限制
limitations- The post-hoc semantic check is explicitly non-canonical and does not make the strict outputs compliant.
- One model; whether stronger models obey the output contract is unknown.
主張
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | supports | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0489 |
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | qualifies | RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3× | ACTIVE | REL-2026-0490 |
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | qualifies | THY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯 | ACTIVE | REL-2026-0491 |
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | supports | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0489 |
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | qualifies | RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3× | ACTIVE | REL-2026-0490 |
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | qualifies | THY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯 | ACTIVE | REL-2026-0491 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | produces | RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | ACTIVE | REL-2026-0488 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/results/RST-2026-0102/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports