ExperimentEXP-2026-0103v0.1
XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03)
The first time a real model was placed inside the IPM instrument: hf.co/empero-ai/Qwythos-9B-v2-GGUF:Q4_K_M served by Ollama on an RTX 3070, run locally on 2026-09-03 by Splice (Claude Code) on Neo.K's authorization through XA-06L — MATH-003, CODE-001, CON-003 × A0–A5 × 2 replicates, 36 trials, 186 invocations, 162 trajectories, telemetry complete on every trial. Sealed REAL_MODEL_PILOT_INCOMPLETE: 33/36 complete, three trials aborted because the same-model verifier returned non-JSON to a strict parser (a real model behaviour, deliberately not re-rolled). SSR = 1.0, SDR = 0.0: scaffolding brought no measured quality gain on these three easy tasks while A5 used 3.10× the device energy and 3.13× the wall time of A0, and the fixed eight-sample conditions A2–A4 used 7.9–9.3×. The apparent A3/A4 quality rise to 0.667 is a missingness artifact; and for two of the three tasks the recorded 'quality' was output-format compliance, not task correctness (post-hoc: 33/33 completed outputs semantically correct; strict output-contract compliance 12/36).
Hypothesis
hypothesis- Experiment A's H1–H4 on the three-task gate matrix: does scaffolding raise quality, at what physical cost, with non-constant marginal yield, and is SSR < 1?
Setup
hardware- NVIDIA GeForce RTX 3070 (8 GiB VRAM; peak 7.59 GiB used, peak 208.9 W, 73 °C), Windows 10 host; physical boundary local-runner-plus-visible-accelerator; energy type device_measured (E-Grade C, CST-B)
software_environment- Ollama serving the model through an OpenAI-compatible endpoint at 127.0.0.1:11434; XA-02/03/04/06/06L v0.1 (hashes verified byte-exact against their manifests before the run); Python 3.14.5; XA-03 collectors system + nvidia_smi at 250 ms target (observed ~335–350 ms)
configurationprovidermode- openai_compatible
provider_id- local-openai-compatible
model_id- hf.co/empero-ai/Qwythos-9B-v2-GGUF:Q4_K_M
base_max_tokens- 1024
temperature- 0.2
top_p- 1.0
seed- 7
budget_control_validated- true
auth_mode- none
matrixtasks- MATH-003
- CODE-001
- CON-003
conditions- A0
- A1
- A2
- A3
- A4
- A5
replicates- 2
allow_code_evaluation- false
telemetrycollectors- system
- nvidia_smi
physical_boundary- local-runner-plus-visible-accelerator
required- false
Procedure
procedure- XA-06L configure → preflight (all checks PASS; isolated MATH-003/A0 quality 1.0) → run (36/36 terminal, 0 runtime failures) → verify gate (3 points xa03_not_complete) → analyze → seal. The three aborted points were not re-run with resume --force: re-rolling until the gate turns green would erase a real failure mode.
Runs
run_count- 1
random_seeds- seed 7 (frozen into every HTTP request); model nondeterminism otherwise uncontrolled
controls- fresh provider per trial; frozen temperature 0.2, top-p 1.0, seed 7; base_max_tokens 1024 with validated 2× budget for A1
- identical initial task text across conditions; calculator tool contract only in A4/A5 generator requests
- scoring after XA-03 finalization; private references never in context; code evaluation disabled on the host
metricsgatestatus- REAL_MODEL_PILOT_INCOMPLETE
verified_complete_count- 33
invalid_or_missingpoint- CODE-001__A3__r2
reason- xa03_not_complete
point- CON-003__A3__r1
reason- xa03_not_complete
point- CON-003__A4__r1
reason- xa03_not_complete
headlinessr- 1.0
sdr- 0.0
scm_device_energy- 3.1019
scm_wall_time- 3.1264
protocol_compliance_rate- 0.3333
quality_available_rate- 0.6111
by_conditionA0quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 10.21
device_energy_j- 1715.6
energy_ratio_vs_A0- 1.0
gpu_peak_memory_gib- 7.23
gpu_memory_residency_gib_s- 72.1
gpu_utilization_integral_s- 6.57
A1quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 9.87
device_energy_j- 1819.0
energy_ratio_vs_A0- 1.06
gpu_peak_memory_gib- 7.32
gpu_memory_residency_gib_s- 70.0
gpu_utilization_integral_s- 6.26
A2quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 79.09
device_energy_j- 15143.0
energy_ratio_vs_A0- 8.827
gpu_peak_memory_gib- 7.28
gpu_memory_residency_gib_s- 570.0
gpu_utilization_integral_s- 54.09
A3quality_mean- 0.6667
quality_n- 3
success_rate- 0.6667
wall_time_s- 99.13
device_energy_j- 15921.5
energy_ratio_vs_A0- 9.281
gpu_peak_memory_gib- 7.34
gpu_memory_residency_gib_s- 714.6
gpu_utilization_integral_s- 70.02
A4quality_mean- 0.6667
quality_n- 3
success_rate- 0.6667
wall_time_s- 69.08
device_energy_j- 13546.2
energy_ratio_vs_A0- 7.896
gpu_peak_memory_gib- 7.28
gpu_memory_residency_gib_s- 496.9
gpu_utilization_integral_s- 46.93
A5quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 31.91
device_energy_j- 5321.5
energy_ratio_vs_A0- 3.102
gpu_peak_memory_gib- 7.24
gpu_memory_residency_gib_s- 228.7
gpu_utilization_integral_s- 22.17
operational_totalsmodel_invocations- 186
trajectories- 162
retries- 0
tool_calls- 0
verifier_passes- 18
candidates_created- 162
candidates_selected- 33
candidates_discarded- 105
failure_events- 6
candidates_abandoned_on_abort- 24
verifier_failurecount- 3
verifier_enabled_trials- 18
rate- 0.16666666666666666
by_conditionA3- 2/6
A4- 1/6
A5- 0/6
strict_protocol_compliancecount- 12
total- 36
rate- 0.3333333333333333
posthoc_semantic_diagnostic (non-canonical)canonical- false
math- 12/12 canonical correct
code- 11/11 selected outputs pass all hidden tests after outer Markdown fence removal
constraint- 10/10 selected outputs satisfy all constraints after format-only normalization
completed_selected_outputs_correct- 33/33
physical_totalstotal_measured_gpu_energy_j- 320800.7795
total_measured_gpu_energy_kwh- 0.0891
summed_trial_wall_time_min- 29.9277
max_gpu_memory_gib- 7.5908
max_gpu_power_w- 208.88
max_gpu_temperature_c- 73.0
mean_sampling_call_wall_fraction- 0.2436
quality_availability_reporting_inconsistencyaggregate_analysis- 22/36
protocol_compliance_csv- 33/36
inconsistency- true
Interpretation
interpretation- As an instrument gate it did its job: real numbers, full telemetry, a sealed and relocatable bundle, and an honest INCOMPLETE. As science it says three things and no more. (1) On three tasks the native single pass already solves, scaffolding cannot show a quality gain — SSR = 1 is a legal null result, and it cost 3.1× (A5) to 9.3× (A3) the device energy of A0; A5 was cheaper than the fixed eight-sample conditions only because its loop stopped early. (2) The instrument's quality axis conflated output-format obedience with task correctness on CODE-001 (correct code inside a Markdown fence) and CON-003 (correct assignment written as A=X, not JSON) — exactly the SyntacticValidity ≠ SemanticCorrectness split Paper 06 predicts, now observed in the lab's own instrument. (3) The same-model verifier's serialization failed in 3 of 18 verifier trials and discarded eight candidates each time; tool access was enabled but never used, so tool and retry effects are unidentified. The A3/A4 'gain' is survivor bias from the aborted low-format trials. Seven instrument revisions are required before XA-07; the dataset stays immutable.
Limitations
limitations- Three easy tasks, two replicates, one 9B model at 4-bit, one machine; not a population-level estimate of anything.
- Gate INCOMPLETE (33/36); CODE-001 quality unmeasured (execution disabled) so 12 of 36 trials have no measured quality; quality-availability is reported inconsistently inside the bundle (22/36 vs 33/36).
- Device-measured GPU energy only — not marginal, not whole-system; telemetry sampling itself cost ~24 % of trial wall time.
- Same-model verifier; no independent or formal verifier condition.
Reproduction
reproduction_instructions- Unpack XA-02/03/04/06/06L as siblings, .\configure.ps1 (local_openai_compatible, base_url http://127.0.0.1:11434, the model id above), .\preflight.ps1, .\run-pilot.ps1, .\seal-results.ps1; verify the sealed bundle against its manifest.json (266 files). The bundle's raw events/telemetry/summaries are unchanged by sealing.
Results
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | produces | RST-2026-0101 Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4 | ACTIVE | REL-2026-0485 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | produces | RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | ACTIVE | REL-2026-0488 |
Recorded fields
completed_at- 2026-09-03
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | extends | EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | ACTIVE | REL-2026-0471 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | extends | EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | ACTIVE | REL-2026-0472 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | uses_benchmark | BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | ACTIVE | REL-2026-0473 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | uses_model | MOD-2026-0006 Qwythos-9B-v2 (Q4_K_M, local, Ollama) | ACTIVE | REL-2026-0474 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | runs_on | SYS-2026-0101 XA-03 — telemetry and run logger (physical execution evidence layer) | ACTIVE | REL-2026-0475 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | runs_on | SYS-2026-0102 XA-04 — A0→A5 model runner and scaffold orchestrator | ACTIVE | REL-2026-0476 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | runs_on | SYS-2026-0103 XA-06 — real-model pilot gate | ACTIVE | REL-2026-0477 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | runs_on | SYS-2026-0104 XA-06L — local real-model execution handoff pack | ACTIVE | REL-2026-0478 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | tests | THY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladder | ACTIVE | REL-2026-0479 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | tests | CLM-2026-0104 F4 — Scaffolding separation | ACTIVE | REL-2026-0480 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | tests | THY-2026-0106 Structured quality, hard gates and the specification–verification separation | ACTIVE | REL-2026-0481 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | tests | THY-2026-0105 Physical computation cost vector and computational spacetime | ACTIVE | REL-2026-0482 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | produced | ART-2026-0119 XA-06 real-model pilot result bundle — Qwythos-9B-v2 Q4_K_M, 36 trials, sealed 2026-09-03 (REAL_MODEL_PILOT_INCOMPLETE) artifact://evemisslab/intelligence-physical-metrology/XA06_REAL_hf.co_empero-ai_Qwythos-9B-v2-GGUF_Q4_K_M_20260903T131913Z.zip | ACTIVE | REL-2026-0483 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | produced | ART-2026-0120 IPM XA-06 first real-model pilot diagnostic v0.1 (2026-09-07) — report, metrics, three figures artifact://evemisslab/intelligence-physical-metrology/IPM_XA06_Qwythos9B_FirstRealPilot_Diagnostic_v0.1.zip | ACTIVE | REL-2026-0484 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | produces | RST-2026-0101 Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4 | ACTIVE | REL-2026-0485 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | produces | RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact | ACTIVE | REL-2026-0488 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0103/
- Machine-readable
/ai/experiments/EXP-2026-0103/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports