EVEMISSLAB

ResultRST-2026-0102v0.1

Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact

Strict output-contract compliance 12/36 (MATH-003 12/12, CODE-001 0/12, CON-003 0/12), but the non-canonical post-hoc check finds every completed selected output correct — math 12/12 canonical correct, code 11/11 selected outputs pass all hidden tests after outer Markdown fence removal, constraint 10/10 selected outputs satisfy all constraints after format-only normalization — 33/33 overall. Verifier-enabled trials 18, verifier-protocol failures 3 (A3 2/6, A4 1/6, A5 0/6), each after eight candidates had been generated. Tool calls 0, retries 0; 24 candidates abandoned on abort are neither selected nor discarded in the schema. Quality availability is reported as 22/36 in the aggregate and 33/36 in protocol_compliance.csv. Telemetry sampling wall fraction 0.244.

Research status
STABLE the current conclusions are relatively stable
Evidence level
E2 Controlled experiment
Result
NEGATIVE
Data basis
REAL MODEL A real language model was executed; the model, its version and its configuration are recorded on the page.
Version
0.1
Updated
2026-09-07
Created
2026-09-03
Domain
Evaluation
Program
PRG-2026-0101 Intelligence Physical Metrology (IPM)
Authors
Neo.K (EveMissLab)
AI collaborators
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT) — 2026-09-07 diagnostic, Splice (Claude Code, Anthropic) — execution and RESULT note

Observed result

metrics
strict_protocol_compliance
count
12
total
36
rate
0.3333333333333333
posthoc_semantic_diagnostic
canonical
false
math
12/12 canonical correct
code
11/11 selected outputs pass all hidden tests after outer Markdown fence removal
constraint
10/10 selected outputs satisfy all constraints after format-only normalization
completed_selected_outputs_correct
33/33
verifier_failure
count
3
verifier_enabled_trials
18
rate
0.16666666666666666
by_condition
A3
2/6
A4
1/6
A5
0/6
operational_totals
model_invocations
186
trajectories
162
retries
0
tool_calls
0
verifier_passes
18
candidates_created
162
candidates_selected
33
candidates_discarded
105
failure_events
6
candidates_abandoned_on_abort
24
quality_availability_reporting
aggregate_analysis
22/36
protocol_compliance_csv
33/36
inconsistency
true
instrumentation
mean_sampling_call_wall_fraction
0.2436
target_sampling_ms
250
observed_cadence_ms
335–350
instrument_revisions_required_before_XA-07
  • R1 typed subject/system failure is valid data, separate from instrument failure
  • R2 split task_semantic_quality / output_contract_compliance / system_completion_reliability / verifier_protocol_reliability
  • R3 failure-aware aggregation (no survivor means)
  • R4 abandoned-candidate accounting
  • R5 preserve verifier parse diagnostics
  • R6 cheaper telemetry (persistent nvidia-smi / NVML)
  • R7 tool-trigger tasks

Interpretation

interpretation
Negative for the instrument, informative for the theory. Task semantic quality ≠ protocol/serialization compliance — the model solved everything it completed and was scored 0.5 on average for not obeying a JSON/no-fence contract; same-model verification reduced system reliability rather than raising quality; tool access ≠ tool utilization (enabled, never used). The bundle stays immutable as XA-06 empirical v0.1; the next step is to revise the instrument and re-run a small validation set, not to re-roll failures until they disappear.

Limitations

limitations
  • The post-hoc semantic check is explicitly non-canonical and does not make the strict outputs compliant.
  • One model; whether stronger models obey the output contract is unknown.

Claims

SourceRelationTargetStatusID
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactsupportsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0489
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactqualifiesRST-2026-0101 Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4ACTIVEREL-2026-0490
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactqualifiesTHY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladderACTIVEREL-2026-0491

Relations

SourceRelationTargetStatusID
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactsupportsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0489
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactqualifiesRST-2026-0101 Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4ACTIVEREL-2026-0490
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactqualifiesTHY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladderACTIVEREL-2026-0491
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03)producesRST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactACTIVEREL-2026-0488

History and provenance

Canonical URL
https://evemisslab.com/ai/results/RST-2026-0102/
Machine-readable
/ai/results/RST-2026-0102/index.json
Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45
Provenance
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports