實驗EXP-2026-0103v0.1
XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03)
第一次把真實模型放進 IPM 儀器:hf.co/empero-ai/Qwythos-9B-v2-GGUF:Q4_K_M 由 Ollama 在 RTX 3070 上服務,2026-09-03 由 Splice(Claude Code)在 Neo.K 授權下透過 XA-06L 於本地執行——MATH-003、CODE-001、CON-003 × A0–A5 × 2 次,36 次試驗、186 次呼叫、162 條軌跡,每次試驗遙測完整。封存為 REAL_MODEL_PILOT_INCOMPLETE:33/36 完成,三次試驗因同模型驗證器對嚴格解析器回了非 JSON 而中止(真實的模型行為,刻意不重擲)。SSR = 1.0、SDR = 0.0:在這三個簡單任務上鷹架沒有帶來可測的品質增益,A5 卻用了 A0 的 3.10× 裝置能量與 3.13× wall time,固定八樣本的 A2–A4 用了 7.9–9.3×。A3/A4 看似升到 0.667 是缺值造成的假象;且三個任務裡有兩個,記錄到的「品質」是輸出格式服從性而非任務正確性(事後檢查:33/33 完成的輸出語意正確;嚴格輸出契約合規 12/36)。
假設
hypothesis- Experiment A's H1–H4 on the three-task gate matrix: does scaffolding raise quality, at what physical cost, with non-constant marginal yield, and is SSR < 1?
設定
benchmark_idshardware- NVIDIA GeForce RTX 3070 (8 GiB VRAM; peak 7.59 GiB used, peak 208.9 W, 73 °C), Windows 10 host; physical boundary local-runner-plus-visible-accelerator; energy type device_measured (E-Grade C, CST-B)
software_environment- Ollama serving the model through an OpenAI-compatible endpoint at 127.0.0.1:11434; XA-02/03/04/06/06L v0.1 (hashes verified byte-exact against their manifests before the run); Python 3.14.5; XA-03 collectors system + nvidia_smi at 250 ms target (observed ~335–350 ms)
configurationprovidermode- openai_compatible
provider_id- local-openai-compatible
model_id- hf.co/empero-ai/Qwythos-9B-v2-GGUF:Q4_K_M
base_max_tokens- 1024
temperature- 0.2
top_p- 1.0
seed- 7
budget_control_validated- true
auth_mode- none
matrixtasks- MATH-003
- CODE-001
- CON-003
conditions- A0
- A1
- A2
- A3
- A4
- A5
replicates- 2
allow_code_evaluation- false
telemetrycollectors- system
- nvidia_smi
physical_boundary- local-runner-plus-visible-accelerator
required- false
程序
procedure- XA-06L configure → preflight (all checks PASS; isolated MATH-003/A0 quality 1.0) → run (36/36 terminal, 0 runtime failures) → verify gate (3 points xa03_not_complete) → analyze → seal. The three aborted points were not re-run with resume --force: re-rolling until the gate turns green would erase a real failure mode.
執行
run_count- 1
random_seeds- seed 7 (frozen into every HTTP request); model nondeterminism otherwise uncontrolled
controls- fresh provider per trial; frozen temperature 0.2, top-p 1.0, seed 7; base_max_tokens 1024 with validated 2× budget for A1
- identical initial task text across conditions; calculator tool contract only in A4/A5 generator requests
- scoring after XA-03 finalization; private references never in context; code evaluation disabled on the host
metricsgatestatus- REAL_MODEL_PILOT_INCOMPLETE
verified_complete_count- 33
invalid_or_missingpoint- CODE-001__A3__r2
reason- xa03_not_complete
point- CON-003__A3__r1
reason- xa03_not_complete
point- CON-003__A4__r1
reason- xa03_not_complete
headlinessr- 1.0
sdr- 0.0
scm_device_energy- 3.1019
scm_wall_time- 3.1264
protocol_compliance_rate- 0.3333
quality_available_rate- 0.6111
by_conditionA0quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 10.21
device_energy_j- 1715.6
energy_ratio_vs_A0- 1.0
gpu_peak_memory_gib- 7.23
gpu_memory_residency_gib_s- 72.1
gpu_utilization_integral_s- 6.57
A1quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 9.87
device_energy_j- 1819.0
energy_ratio_vs_A0- 1.06
gpu_peak_memory_gib- 7.32
gpu_memory_residency_gib_s- 70.0
gpu_utilization_integral_s- 6.26
A2quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 79.09
device_energy_j- 15143.0
energy_ratio_vs_A0- 8.827
gpu_peak_memory_gib- 7.28
gpu_memory_residency_gib_s- 570.0
gpu_utilization_integral_s- 54.09
A3quality_mean- 0.6667
quality_n- 3
success_rate- 0.6667
wall_time_s- 99.13
device_energy_j- 15921.5
energy_ratio_vs_A0- 9.281
gpu_peak_memory_gib- 7.34
gpu_memory_residency_gib_s- 714.6
gpu_utilization_integral_s- 70.02
A4quality_mean- 0.6667
quality_n- 3
success_rate- 0.6667
wall_time_s- 69.08
device_energy_j- 13546.2
energy_ratio_vs_A0- 7.896
gpu_peak_memory_gib- 7.28
gpu_memory_residency_gib_s- 496.9
gpu_utilization_integral_s- 46.93
A5quality_mean- 0.5
quality_n- 4
success_rate- 0.5
wall_time_s- 31.91
device_energy_j- 5321.5
energy_ratio_vs_A0- 3.102
gpu_peak_memory_gib- 7.24
gpu_memory_residency_gib_s- 228.7
gpu_utilization_integral_s- 22.17
operational_totalsmodel_invocations- 186
trajectories- 162
retries- 0
tool_calls- 0
verifier_passes- 18
candidates_created- 162
candidates_selected- 33
candidates_discarded- 105
failure_events- 6
candidates_abandoned_on_abort- 24
verifier_failurecount- 3
verifier_enabled_trials- 18
rate- 0.16666666666666666
by_conditionA3- 2/6
A4- 1/6
A5- 0/6
strict_protocol_compliancecount- 12
total- 36
rate- 0.3333333333333333
posthoc_semantic_diagnostic (non-canonical)canonical- false
math- 12/12 canonical correct
code- 11/11 selected outputs pass all hidden tests after outer Markdown fence removal
constraint- 10/10 selected outputs satisfy all constraints after format-only normalization
completed_selected_outputs_correct- 33/33
physical_totalstotal_measured_gpu_energy_j- 320800.7795
total_measured_gpu_energy_kwh- 0.0891
summed_trial_wall_time_min- 29.9277
max_gpu_memory_gib- 7.5908
max_gpu_power_w- 208.88
max_gpu_temperature_c- 73.0
mean_sampling_call_wall_fraction- 0.2436
quality_availability_reporting_inconsistencyaggregate_analysis- 22/36
protocol_compliance_csv- 33/36
inconsistency- true
詮釋
interpretation- As an instrument gate it did its job: real numbers, full telemetry, a sealed and relocatable bundle, and an honest INCOMPLETE. As science it says three things and no more. (1) On three tasks the native single pass already solves, scaffolding cannot show a quality gain — SSR = 1 is a legal null result, and it cost 3.1× (A5) to 9.3× (A3) the device energy of A0; A5 was cheaper than the fixed eight-sample conditions only because its loop stopped early. (2) The instrument's quality axis conflated output-format obedience with task correctness on CODE-001 (correct code inside a Markdown fence) and CON-003 (correct assignment written as A=X, not JSON) — exactly the SyntacticValidity ≠ SemanticCorrectness split Paper 06 predicts, now observed in the lab's own instrument. (3) The same-model verifier's serialization failed in 3 of 18 verifier trials and discarded eight candidates each time; tool access was enabled but never used, so tool and retry effects are unidentified. The A3/A4 'gain' is survivor bias from the aborted low-format trials. Seven instrument revisions are required before XA-07; the dataset stays immutable.
限制
limitations- Three easy tasks, two replicates, one 9B model at 4-bit, one machine; not a population-level estimate of anything.
- Gate INCOMPLETE (33/36); CODE-001 quality unmeasured (execution disabled) so 12 of 36 trials have no measured quality; quality-availability is reported inconsistently inside the bundle (22/36 vs 33/36).
- Device-measured GPU energy only — not marginal, not whole-system; telemetry sampling itself cost ~24 % of trial wall time.
- Same-model verifier; no independent or formal verifier condition.
重現
reproduction_instructions- Unpack XA-02/03/04/06/06L as siblings, .\configure.ps1 (local_openai_compatible, base_url http://127.0.0.1:11434, the model id above), .\preflight.ps1, .\run-pilot.ps1, .\seal-results.ps1; verify the sealed bundle against its manifest.json (266 files). The bundle's raw events/telemetry/summaries are unchanged by sealing.
結果
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | produces | RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3× | ACTIVE | REL-2026-0485 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | produces | RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | ACTIVE | REL-2026-0488 |
記錄欄位
completed_at- 2026-09-03
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | extends | EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | ACTIVE | REL-2026-0471 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | extends | EXP-2026-0102 XA-05——以腳本化供應商跑的 36 試驗端到端煙霧閘(合成) | ACTIVE | REL-2026-0472 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | uses_benchmark | BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包 | ACTIVE | REL-2026-0473 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | uses_model | MOD-2026-0006 Qwythos-9B-v2(Q4_K_M,本地,Ollama) | ACTIVE | REL-2026-0474 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | runs_on | SYS-2026-0101 XA-03——遙測與執行記錄器(物理執行證據層) | ACTIVE | REL-2026-0475 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | runs_on | SYS-2026-0102 XA-04——A0→A5 模型執行器與鷹架調度器 | ACTIVE | REL-2026-0476 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | runs_on | SYS-2026-0103 XA-06——真實模型 pilot 閘 | ACTIVE | REL-2026-0477 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | runs_on | SYS-2026-0104 XA-06L——本地真實模型執行交接包 | ACTIVE | REL-2026-0478 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | tests | THY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯 | ACTIVE | REL-2026-0479 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | tests | CLM-2026-0104 F4——鷹架分離 | ACTIVE | REL-2026-0480 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | tests | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0481 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | tests | THY-2026-0105 物理計算成本向量與計算時空 | ACTIVE | REL-2026-0482 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | produced | ART-2026-0119 XA-06 real-model pilot result bundle — Qwythos-9B-v2 Q4_K_M, 36 trials, sealed 2026-09-03 (REAL_MODEL_PILOT_INCOMPLETE) artifact://evemisslab/intelligence-physical-metrology/XA06_REAL_hf.co_empero-ai_Qwythos-9B-v2-GGUF_Q4_K_M_20260903T131913Z.zip | ACTIVE | REL-2026-0483 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | produced | ART-2026-0120 IPM XA-06 first real-model pilot diagnostic v0.1 (2026-09-07) — report, metrics, three figures artifact://evemisslab/intelligence-physical-metrology/IPM_XA06_Qwythos9B_FirstRealPilot_Diagnostic_v0.1.zip | ACTIVE | REL-2026-0484 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | produces | RST-2026-0101 三個簡單任務上的鷹架響應:SSR = 1.0,A5 的裝置能量 3.1×、A2–A4 7.9–9.3× | ACTIVE | REL-2026-0485 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | produces | RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | ACTIVE | REL-2026-0488 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0103/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports