ExperimentEXP-2026-0101v0.1
Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)
The first v0.2 experiment: for one model and a fixed task set, climb the scaffold ladder A0 native single pass → A1 extended trajectory → A2 multi-sample → A3 verifier → A4 deterministic local tools → A5 bounded full agentic loop, recording quality and physical cost at every level to obtain a scaffolding response curve, SSR/SDR, scaffold cost multipliers and marginal yields. Five hypotheses (H1 gain exists, H2 gain has physical cost, H3 marginal yield is non-constant, H4 native and system capability are distinguishable, H5 models have different scaffolding profiles), pre-registered quality projections, initial-information equality across A0–A3, budget caps, n = 5 pilot / 20 formal replicates, failure classification and the rule that a null result is not an experiment failure. Status READY FOR PILOT; the thirty-task run itself has not been executed — only the three-task instrument gates below.
Hypothesis
hypothesis- H1 Q_A5 > Q_A0 for at least some non-trivial tasks; H2 E, V_C, T rise with it; H3 marginal yield differs by stage; H4 SSR < 1 stably; H5 SSR and SCM differ across models even at equal Q_A5.
Setup
software_environment- protocol document + JSON run schema + YAML example run (EML-IPM-XA-01 v0.1)
Procedure
procedure- 30 tasks (10 math, 10 code, 10 constraint; Easy/Medium/Hard predefined) × A0–A5 × n replicates; A2 8 trajectories with deterministic majority; A3 typed verifier (same-model / independent / formal); A4 ≤ 4 deterministic local tool calls; A5 ≤ 16 invocations, ≤ 8 tool calls, ≤ 3 retry cycles with explicit termination; seeds and decoding frozen; warm weights, clean task state; run IDs IPM-XA-{model}-{task}-{condition}-{replicate}; paired within-task statistics with bootstrap CIs and effect sizes.
Runs
run_count- 0
controls- identical prompt, initial context, decoding, system instruction and model version across conditions
- initial-information equality A0–A3; external information gain marked for A4/A5
- no cross-condition leakage; budget self-extension forbidden
metricsplanned outputs- ΔQ_k = Q_k − Q_0
- SSR = Q_0 / Q_5, SDR = 1 − SSR
- SCM_j = C_5,j / C_0,j per cost axis
- marginal yield Y_k,j
- response curves Q vs T, E, invocations, device-time
- brute-force flag ΔQ < 0.01 with ΔC/C > 0.5 (exploratory thresholds)
- selection waste ratio, discarded work
minimum physical telemetry- T_wall
- E_device (E-Grade C)
- M_peak
- V_C
status- READY FOR PILOT
Interpretation
interpretation- A protocol, not a result. Its instrument (XA-02…XA-06L) was validated synthetically and then used once on three easy tasks with a real local model; whether a scaffolding response curve exists on non-trivial tasks is still open.
Limitations
limitations- Deliberately does not attempt μI identification, lifecycle energy, cross-substrate comparison, high-ambiguity quality, full Shapley attribution or multi-agent settings.
- Pilot budgets are reference values, not IPM standards.
Reproduction
reproduction_instructions- Implement the ladder with XA-04 against XA-02 tasks, log with XA-03, run through XA-06/XA-06L; see EXP-2026-0103 for the first real execution.
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | uses_benchmark | BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | ACTIVE | REL-2026-0461 |
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | runs_on | SYS-2026-0102 XA-04 — A0→A5 model runner and scaffold orchestrator | ACTIVE | REL-2026-0462 |
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | tests | THY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladder | ACTIVE | REL-2026-0463 |
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | tests | CLM-2026-0104 F4 — Scaffolding separation | ACTIVE | REL-2026-0464 |
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | produced | ART-2026-0112 EML-IPM-XA-01 v0.1 — Experiment A protocol, run schema and example run artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_Experiment_A_Protocol_Package_v0.1.zip | ACTIVE | REL-2026-0465 |
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | extends | EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | ACTIVE | REL-2026-0466 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | extends | EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | ACTIVE | REL-2026-0471 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0101/
- Machine-readable
/ai/experiments/EXP-2026-0101/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports