實驗EXP-2026-0101v0.1
Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1
第一個 v0.2 實驗:對同一模型與固定任務集,沿鷹架階梯 A0 原生單次 → A1 放大軌跡 → A2 多樣本 → A3 驗證器 → A4 確定性本地工具 → A5 有界完整 agentic 迴圈往上爬,每一級同時記品質與物理成本,得出鷹架響應曲線、SSR/SDR、鷹架成本倍率與邊際產率。五個假說(H1 增益存在、H2 增益有物理成本、H3 邊際產率非常數、H4 原生與系統能力可區分、H5 不同模型有不同鷹架剖面)、預先登記的品質投影、A0–A3 初始資訊相等、預算上限、pilot n = 5/正式 n = 20 次重複、失敗分類,以及「null 結果不是實驗失敗」的規則。狀態 READY FOR PILOT;三十題的正式執行尚未進行——只跑過下面的三題儀器閘。
假設
hypothesis- H1 Q_A5 > Q_A0 for at least some non-trivial tasks; H2 E, V_C, T rise with it; H3 marginal yield differs by stage; H4 SSR < 1 stably; H5 SSR and SCM differ across models even at equal Q_A5.
設定
benchmark_idssoftware_environment- protocol document + JSON run schema + YAML example run (EML-IPM-XA-01 v0.1)
程序
procedure- 30 tasks (10 math, 10 code, 10 constraint; Easy/Medium/Hard predefined) × A0–A5 × n replicates; A2 8 trajectories with deterministic majority; A3 typed verifier (same-model / independent / formal); A4 ≤ 4 deterministic local tool calls; A5 ≤ 16 invocations, ≤ 8 tool calls, ≤ 3 retry cycles with explicit termination; seeds and decoding frozen; warm weights, clean task state; run IDs IPM-XA-{model}-{task}-{condition}-{replicate}; paired within-task statistics with bootstrap CIs and effect sizes.
執行
run_count- 0
controls- identical prompt, initial context, decoding, system instruction and model version across conditions
- initial-information equality A0–A3; external information gain marked for A4/A5
- no cross-condition leakage; budget self-extension forbidden
metricsplanned outputs- ΔQ_k = Q_k − Q_0
- SSR = Q_0 / Q_5, SDR = 1 − SSR
- SCM_j = C_5,j / C_0,j per cost axis
- marginal yield Y_k,j
- response curves Q vs T, E, invocations, device-time
- brute-force flag ΔQ < 0.01 with ΔC/C > 0.5 (exploratory thresholds)
- selection waste ratio, discarded work
minimum physical telemetry- T_wall
- E_device (E-Grade C)
- M_peak
- V_C
status- READY FOR PILOT
詮釋
interpretation- A protocol, not a result. Its instrument (XA-02…XA-06L) was validated synthetically and then used once on three easy tasks with a real local model; whether a scaffolding response curve exists on non-trivial tasks is still open.
限制
limitations- Deliberately does not attempt μI identification, lifecycle energy, cross-substrate comparison, high-ambiguity quality, full Shapley attribution or multi-agent settings.
- Pilot budgets are reference values, not IPM standards.
重現
reproduction_instructions- Implement the ladder with XA-04 against XA-02 tasks, log with XA-03, run through XA-06/XA-06L; see EXP-2026-0103 for the first real execution.
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | uses_benchmark | BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包 | ACTIVE | REL-2026-0461 |
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | runs_on | SYS-2026-0102 XA-04——A0→A5 模型執行器與鷹架調度器 | ACTIVE | REL-2026-0462 |
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | tests | THY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯 | ACTIVE | REL-2026-0463 |
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | tests | CLM-2026-0104 F4——鷹架分離 | ACTIVE | REL-2026-0464 |
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | produced | ART-2026-0112 EML-IPM-XA-01 v0.1 — Experiment A protocol, run schema and example run artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_Experiment_A_Protocol_Package_v0.1.zip | ACTIVE | REL-2026-0465 |
EXP-2026-0102 XA-05——以腳本化供應商跑的 36 試驗端到端煙霧閘(合成) | extends | EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | ACTIVE | REL-2026-0466 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | extends | EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | ACTIVE | REL-2026-0471 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0101/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports