EVEMISSLAB
English

計畫PRG-2026-0101v0.1

智能的物理計量(IPM)

這個研究計畫問的不是「AI 有幾分聰明」,而是一個答案讓物理系統付出了什麼:花了幾次使用者回合、幾條生成軌跡、幾層隱藏 LOOP,做了多少有效語意工作,占用多少能量與計算時空,依賴多少外部鷹架,最後換回多少可驗證品質。十篇理論論文(2026-09-02)定義了測量物件——任務、品質、語意工作、物理計算、鷹架能力、測量詮釋資料——與五個可證偽命題;v0.2 實驗協定(Experiment A:single-pass vs scaffolded)落實為 XA-02…XA-06L 儀器套件,並在本地 9B 模型上執行過一次。

研究狀態
ACTIVE 正在持續研究
證據等級
E2 受控實驗
版本
0.1
更新
2026-09-07
建立
2026-09-02
領域
Evaluation, Computation, Cognitive Science, AI Architecture
作者
Neo.K (EveMissLab)
AI 協作
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT)

目標

goals
  • Replace token, FLOP, benchmark score and 'one turn' as units of intelligence with a typed event 𝔍_IPM = (task, quality, semantic work, physical computation, scaffolding, metadata).
  • Measure the relation physical computation → effective semantic work → verifiable quality, and compare systems on a Pareto frontier instead of a single score.
  • Put the five falsifiable propositions — token hypothesis, FLOPs sufficiency, binary burden, scaffolding separation, semantic intermediate utility — in front of experiments in the order A → D → B → C → E.
  • Publish under the IPM Minimum Reporting Standard: boundary, energy type, hidden work, uncertainty and measurement grade on every number.

開放問題

open_questions
  • μI has no operational identification yet (Experiment C); the series itself says μI must pass a predictive/explanatory utility test or be revised or eliminated.
  • The pilot's quality axis measured output-format compliance for two of three tasks; task semantic quality, output-contract compliance, system completion reliability and verifier reliability have to be separated before scaling.
  • Same-model verifier serialization failed in 3 of 18 verifier trials; candidates abandoned on abort are not yet accounted; telemetry sampling costs about 24 % of trial wall time.
  • CODE tasks have no measured quality until the pilot runs inside a disposable sandbox with code evaluation enabled.
  • Experiments B, C, D and E are declared, not run.

里程碑

milestones
  • 2026-09-02: EML-IPM v0.1 theoretical series complete, 10/10 papers with a SHA-256 manifest; canonical index with unified notation and the v0.2 experimental entry point.
  • 2026-09-02: Experiment A protocol (EML-IPM-XA-01 v0.1, READY FOR PILOT).
  • 2026-09-03: instrument packages XA-02 (30-task pack), XA-03 (telemetry logger), XA-04 (A0→A5 runner), XA-05 (36-trial synthetic smoke gate), XA-06 (real-model pilot gate), XA-06L (local execution handoff).
  • 2026-09-03: first real-model pilot on a local Qwythos-9B-v2 — 36 trials, sealed REAL_MODEL_PILOT_INCOMPLETE (33/36 complete).
  • 2026-09-07: diagnostic analysis of the pilot with seven instrument revisions required before XA-07.

此計畫下的研究線與系統

來源關係目標狀態ID
RES-2026-0101 執行與物理線——一個答案在回合、語意工作、能量與計算時空上的代價belongs_toPRG-2026-0101 智能的物理計量(IPM)ACTIVEREL-2026-0353
RES-2026-0102 品質線——不叫人打分數也能量成果品質belongs_toPRG-2026-0101 智能的物理計量(IPM)ACTIVEREL-2026-0354
RES-2026-0103 能力線——拿掉 LOOP 還剩多少智能,以及統一的智能事件belongs_toPRG-2026-0101 智能的物理計量(IPM)ACTIVEREL-2026-0355

關係

來源關係目標狀態ID
PRG-2026-0101 智能的物理計量(IPM)released_asART-2026-0101 EML-IPM v0.1 canonical series package — 10 papers + canonical index + SHA-256 manifest (UTF-8 Markdown) artifact://evemisslab/intelligence-physical-metrology/IPM_v0.1_Canonical_Series_Package.zipACTIVEREL-2026-0390
RES-2026-0101 執行與物理線——一個答案在回合、語意工作、能量與計算時空上的代價belongs_toPRG-2026-0101 智能的物理計量(IPM)ACTIVEREL-2026-0353
RES-2026-0102 品質線——不叫人打分數也能量成果品質belongs_toPRG-2026-0101 智能的物理計量(IPM)ACTIVEREL-2026-0354
RES-2026-0103 能力線——拿掉 LOOP 還剩多少智能,以及統一的智能事件belongs_toPRG-2026-0101 智能的物理計量(IPM)ACTIVEREL-2026-0355

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/programs/PRG-2026-0101/
機器可讀
/ai/programs/PRG-2026-0101/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports