EVEMISSLAB

ExperimentEXP-2026-0101v0.1

Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)

The first v0.2 experiment: for one model and a fixed task set, climb the scaffold ladder A0 native single pass → A1 extended trajectory → A2 multi-sample → A3 verifier → A4 deterministic local tools → A5 bounded full agentic loop, recording quality and physical cost at every level to obtain a scaffolding response curve, SSR/SDR, scaffold cost multipliers and marginal yields. Five hypotheses (H1 gain exists, H2 gain has physical cost, H3 marginal yield is non-constant, H4 native and system capability are distinguishable, H5 models have different scaffolding profiles), pre-registered quality projections, initial-information equality across A0–A3, budget caps, n = 5 pilot / 20 formal replicates, failure classification and the rule that a null result is not an experiment failure. Status READY FOR PILOT; the thirty-task run itself has not been executed — only the three-task instrument gates below.

Research status
ACTIVE under continuous research
Evidence level
E0 Concept only
Result
INCONCLUSIVE
Data basis
NOT RUN Designed and validated as a harness; never executed with a real model.
Version
0.1
Updated
2026-09-02
Created
2026-09-02
Domain
Evaluation, Agent Systems
Program
PRG-2026-0101 Intelligence Physical Metrology (IPM)
Authors
Neo.K (EveMissLab)
AI collaborators
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT)

Hypothesis

hypothesis
H1 Q_A5 > Q_A0 for at least some non-trivial tasks; H2 E, V_C, T rise with it; H3 marginal yield differs by stage; H4 SSR < 1 stably; H5 SSR and SCM differ across models even at equal Q_A5.

Setup

benchmark_ids
software_environment
protocol document + JSON run schema + YAML example run (EML-IPM-XA-01 v0.1)

Procedure

procedure
30 tasks (10 math, 10 code, 10 constraint; Easy/Medium/Hard predefined) × A0–A5 × n replicates; A2 8 trajectories with deterministic majority; A3 typed verifier (same-model / independent / formal); A4 ≤ 4 deterministic local tool calls; A5 ≤ 16 invocations, ≤ 8 tool calls, ≤ 3 retry cycles with explicit termination; seeds and decoding frozen; warm weights, clean task state; run IDs IPM-XA-{model}-{task}-{condition}-{replicate}; paired within-task statistics with bootstrap CIs and effect sizes.

Runs

run_count
0
controls
  • identical prompt, initial context, decoding, system instruction and model version across conditions
  • initial-information equality A0–A3; external information gain marked for A4/A5
  • no cross-condition leakage; budget self-extension forbidden
metrics
planned outputs
  • ΔQ_k = Q_k − Q_0
  • SSR = Q_0 / Q_5, SDR = 1 − SSR
  • SCM_j = C_5,j / C_0,j per cost axis
  • marginal yield Y_k,j
  • response curves Q vs T, E, invocations, device-time
  • brute-force flag ΔQ < 0.01 with ΔC/C > 0.5 (exploratory thresholds)
  • selection waste ratio, discarded work
minimum physical telemetry
  • T_wall
  • E_device (E-Grade C)
  • M_peak
  • V_C
status
READY FOR PILOT

Interpretation

interpretation
A protocol, not a result. Its instrument (XA-02…XA-06L) was validated synthetically and then used once on three easy tasks with a real local model; whether a scaffolding response curve exists on non-trivial tasks is still open.

Limitations

limitations
  • Deliberately does not attempt μI identification, lifecycle energy, cross-substrate comparison, high-ambiguity quality, full Shapley attribution or multi-agent settings.
  • Pilot budgets are reference values, not IPM standards.

Reproduction

reproduction_instructions
Implement the ladder with XA-04 against XA-02 tasks, log with XA-03, run through XA-06/XA-06L; see EXP-2026-0103 for the first real execution.

Relations

SourceRelationTargetStatusID
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)uses_benchmarkBEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding responseACTIVEREL-2026-0461
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)runs_onSYS-2026-0102 XA-04 — A0→A5 model runner and scaffold orchestratorACTIVEREL-2026-0462
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)testsTHY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladderACTIVEREL-2026-0463
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)testsCLM-2026-0104 F4 — Scaffolding separationACTIVEREL-2026-0464
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)producedART-2026-0112 EML-IPM-XA-01 v0.1 — Experiment A protocol, run schema and example run artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_Experiment_A_Protocol_Package_v0.1.zipACTIVEREL-2026-0465
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic)extendsEXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)ACTIVEREL-2026-0466
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03)extendsEXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)ACTIVEREL-2026-0471

History and provenance

Canonical URL
https://evemisslab.com/ai/experiments/EXP-2026-0101/
Machine-readable
/ai/experiments/EXP-2026-0101/index.json
Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45
Provenance
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports