ExperimentEXP-2026-0102v0.1
XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic)
Runs MATH-003, CODE-001 and CON-003 × A0–A5 × 2 replicates through XA-04 + XA-03 + XA-02 with a ScriptedProvider whose outputs were constructed so the expected quality curve is known in advance (A0 0.30, A1 0.633, A2–A5 1.0 → SSR 0.30 / SDR 0.70 as fixtures). 36 trials, 180 invocations, 168 trajectories, 6 retries, 6 tool calls, 24 verifier passes; candidates 168 = 36 selected + 132 discarded; 0 accounting mismatches, 0 private-reference leaks; energy null on purpose (NullCollector) to prove Unknown ≠ 0. The package states scientific_interpretation_allowed: false.
Hypothesis
hypothesis- Engineering only: the XA-01→XA-04 pipeline reconstructs a pre-designed A0→A5 quality structure with exact candidate accounting and no private-reference leakage.
Setup
software_environment- XA-05 v0.1 over XA-02/03/04 (hashes declared in README and asserted by this extractor)
Procedure
procedure- python -m pytest under IPM_XA02_ROOT / IPM_XA03_ROOT / IPM_XA04_ROOT; verify_canonical_output('canonical_smoke_run') re-validates the packaged run from the extracted package.
Runs
run_count- 1
controls- scripted outputs with a known oracle
- NullCollector so no physical number can masquerade as measurement
- dependency packages pinned by SHA-256
metricsgatetrial_count- 36
all_trials_complete- true
accounting_mismatches- 0
private_sentinel_leaks- 0
ssr- 0.3
sdr- 0.7
energy_scmscientific_interpretation_allowed- false
execution_totalscreated- 168
discarded- 132
model- 180
retry- 6
selected- 36
tool- 6
trajectory- 168
verifier- 24
Interpretation
interpretation- The instrument works as an accounting machine: candidate identity 168 = 36 + 132 holds, nothing leaks, the designed curve comes back out. Nothing here is a statement about any model — the package forbids reading it that way, and this record carries the SYNTHETIC badge for the same reason.
Limitations
limitations- Synthetic fixtures; three tasks; no physical telemetry by design.
Reproduction
reproduction_instructions- Extract XA-02, XA-03, XA-04 and XA-05 as siblings, export the three roots, PYTHONPATH=. python -m pytest -q; or verify the shipped canonical_smoke_run/.
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | extends | EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | ACTIVE | REL-2026-0466 |
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | uses_benchmark | BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | ACTIVE | REL-2026-0467 |
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | runs_on | SYS-2026-0101 XA-03 — telemetry and run logger (physical execution evidence layer) | ACTIVE | REL-2026-0468 |
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | runs_on | SYS-2026-0102 XA-04 — A0→A5 model runner and scaffold orchestrator | ACTIVE | REL-2026-0469 |
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | produced | ART-2026-0116 EML-IPM-XA-05 v0.1 — 36-trial end-to-end synthetic smoke gate with canonical scripted run artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_XA05_36Trial_SmokeGate_v0.1.zip | ACTIVE | REL-2026-0470 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | extends | EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | ACTIVE | REL-2026-0472 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0102/
- Machine-readable
/ai/experiments/EXP-2026-0102/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports