EVEMISSLAB
English

實驗EXP-2026-0101v0.1

Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1

第一個 v0.2 實驗:對同一模型與固定任務集,沿鷹架階梯 A0 原生單次 → A1 放大軌跡 → A2 多樣本 → A3 驗證器 → A4 確定性本地工具 → A5 有界完整 agentic 迴圈往上爬,每一級同時記品質與物理成本,得出鷹架響應曲線、SSR/SDR、鷹架成本倍率與邊際產率。五個假說(H1 增益存在、H2 增益有物理成本、H3 邊際產率非常數、H4 原生與系統能力可區分、H5 不同模型有不同鷹架剖面)、預先登記的品質投影、A0–A3 初始資訊相等、預算上限、pilot n = 5/正式 n = 20 次重複、失敗分類,以及「null 結果不是實驗失敗」的規則。狀態 READY FOR PILOT;三十題的正式執行尚未進行——只跑過下面的三題儀器閘。

研究狀態
ACTIVE 正在持續研究
證據等級
E0 僅有概念
結果
INCONCLUSIVE
資料基礎
NOT RUN 只設計並驗證了 harness,從未用真實模型執行。
版本
0.1
更新
2026-09-02
建立
2026-09-02
領域
Evaluation, Agent Systems
計畫
PRG-2026-0101 智能的物理計量(IPM)
作者
Neo.K (EveMissLab)
AI 協作
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT)

假設

hypothesis
H1 Q_A5 > Q_A0 for at least some non-trivial tasks; H2 E, V_C, T rise with it; H3 marginal yield differs by stage; H4 SSR < 1 stably; H5 SSR and SCM differ across models even at equal Q_A5.

設定

benchmark_ids
software_environment
protocol document + JSON run schema + YAML example run (EML-IPM-XA-01 v0.1)

程序

procedure
30 tasks (10 math, 10 code, 10 constraint; Easy/Medium/Hard predefined) × A0–A5 × n replicates; A2 8 trajectories with deterministic majority; A3 typed verifier (same-model / independent / formal); A4 ≤ 4 deterministic local tool calls; A5 ≤ 16 invocations, ≤ 8 tool calls, ≤ 3 retry cycles with explicit termination; seeds and decoding frozen; warm weights, clean task state; run IDs IPM-XA-{model}-{task}-{condition}-{replicate}; paired within-task statistics with bootstrap CIs and effect sizes.

執行

run_count
0
controls
  • identical prompt, initial context, decoding, system instruction and model version across conditions
  • initial-information equality A0–A3; external information gain marked for A4/A5
  • no cross-condition leakage; budget self-extension forbidden
metrics
planned outputs
  • ΔQ_k = Q_k − Q_0
  • SSR = Q_0 / Q_5, SDR = 1 − SSR
  • SCM_j = C_5,j / C_0,j per cost axis
  • marginal yield Y_k,j
  • response curves Q vs T, E, invocations, device-time
  • brute-force flag ΔQ < 0.01 with ΔC/C > 0.5 (exploratory thresholds)
  • selection waste ratio, discarded work
minimum physical telemetry
  • T_wall
  • E_device (E-Grade C)
  • M_peak
  • V_C
status
READY FOR PILOT

詮釋

interpretation
A protocol, not a result. Its instrument (XA-02…XA-06L) was validated synthetically and then used once on three easy tasks with a real local model; whether a scaffolding response curve exists on non-trivial tasks is still open.

限制

limitations
  • Deliberately does not attempt μI identification, lifecycle energy, cross-substrate comparison, high-ambiguity quality, full Shapley attribution or multi-agent settings.
  • Pilot budgets are reference values, not IPM standards.

重現

reproduction_instructions
Implement the ladder with XA-04 against XA-02 tasks, log with XA-03, run through XA-06/XA-06L; see EXP-2026-0103 for the first real execution.

關係

來源關係目標狀態ID
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1uses_benchmarkBEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包ACTIVEREL-2026-0461
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1runs_onSYS-2026-0102 XA-04——A0→A5 模型執行器與鷹架調度器ACTIVEREL-2026-0462
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1testsTHY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯ACTIVEREL-2026-0463
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1testsCLM-2026-0104 F4——鷹架分離ACTIVEREL-2026-0464
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1producedART-2026-0112 EML-IPM-XA-01 v0.1 — Experiment A protocol, run schema and example run artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_Experiment_A_Protocol_Package_v0.1.zipACTIVEREL-2026-0465
EXP-2026-0102 XA-05——以腳本化供應商跑的 36 試驗端到端煙霧閘(合成)extendsEXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1ACTIVEREL-2026-0466
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03)extendsEXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1ACTIVEREL-2026-0471

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/experiments/EXP-2026-0101/
機器可讀
/ai/experiments/EXP-2026-0101/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports