ResultRST-2026-0102v0.1
Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact
Strict output-contract compliance 12/36 (MATH-003 12/12, CODE-001 0/12, CON-003 0/12), but the non-canonical post-hoc check finds every completed selected output correct — math 12/12 canonical correct, code 11/11 selected outputs pass all hidden tests after outer Markdown fence removal, constraint 10/10 selected outputs satisfy all constraints after format-only normalization — 33/33 overall. Verifier-enabled trials 18, verifier-protocol failures 3 (A3 2/6, A4 1/6, A5 0/6), each after eight candidates had been generated. Tool calls 0, retries 0; 24 candidates abandoned on abort are neither selected nor discarded in the schema. Quality availability is reported as 22/36 in the aggregate and 33/36 in protocol_compliance.csv. Telemetry sampling wall fraction 0.244.
Observed result
metricsstrict_protocol_compliancecount- 12
total- 36
rate- 0.3333333333333333
posthoc_semantic_diagnosticcanonical- false
math- 12/12 canonical correct
code- 11/11 selected outputs pass all hidden tests after outer Markdown fence removal
constraint- 10/10 selected outputs satisfy all constraints after format-only normalization
completed_selected_outputs_correct- 33/33
verifier_failurecount- 3
verifier_enabled_trials- 18
rate- 0.16666666666666666
by_conditionA3- 2/6
A4- 1/6
A5- 0/6
operational_totalsmodel_invocations- 186
trajectories- 162
retries- 0
tool_calls- 0
verifier_passes- 18
candidates_created- 162
candidates_selected- 33
candidates_discarded- 105
failure_events- 6
candidates_abandoned_on_abort- 24
quality_availability_reportingaggregate_analysis- 22/36
protocol_compliance_csv- 33/36
inconsistency- true
instrumentationmean_sampling_call_wall_fraction- 0.2436
target_sampling_ms- 250
observed_cadence_ms- 335–350
instrument_revisions_required_before_XA-07- R1 typed subject/system failure is valid data, separate from instrument failure
- R2 split task_semantic_quality / output_contract_compliance / system_completion_reliability / verifier_protocol_reliability
- R3 failure-aware aggregation (no survivor means)
- R4 abandoned-candidate accounting
- R5 preserve verifier parse diagnostics
- R6 cheaper telemetry (persistent nvidia-smi / NVML)
- R7 tool-trigger tasks
Interpretation
interpretation- Negative for the instrument, informative for the theory. Task semantic quality ≠ protocol/serialization compliance — the model solved everything it completed and was scored 0.5 on average for not obeying a JSON/no-fence contract; same-model verification reduced system reliability rather than raising quality; tool access ≠ tool utilization (enabled, never used). The bundle stays immutable as XA-06 empirical v0.1; the next step is to revise the instrument and re-run a small validation set, not to re-roll failures until they disappear.
Limitations
limitations- The post-hoc semantic check is explicitly non-canonical and does not make the strict outputs compliant.
- One model; whether stronger models obey the output contract is unknown.
Claims
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | supports | THY-2026-0106 Structured quality, hard gates and the specification–verification separation | ACTIVE | REL-2026-0489 |
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | qualifies | RST-2026-0101 Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4 | ACTIVE | REL-2026-0490 |
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | qualifies | THY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladder | ACTIVE | REL-2026-0491 |
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | supports | THY-2026-0106 Structured quality, hard gates and the specification–verification separation | ACTIVE | REL-2026-0489 |
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | qualifies | RST-2026-0101 Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4 | ACTIVE | REL-2026-0490 |
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | qualifies | THY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladder | ACTIVE | REL-2026-0491 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | produces | RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | ACTIVE | REL-2026-0488 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/results/RST-2026-0102/
- Machine-readable
/ai/results/RST-2026-0102/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports