ProgramPRG-2026-0101v0.1
Intelligence Physical Metrology (IPM)
A research program that asks not how smart an AI is but what one answer costs a physical system: how many user turns, generation trajectories and hidden loops it took, how much effective semantic work was done, how much energy and computational spacetime was occupied, how much external scaffolding was leaned on, and how much verifiable quality came back. Ten theoretical papers (2026-09-02) define the measurement objects — task, quality, semantic work, physical computation, scaffolding capability, measurement metadata — and five falsifiable propositions; the v0.2 experimental protocol (Experiment A, single-pass vs scaffolded) is instantiated as the XA-02…XA-06L instrument packages and was executed once on a local 9B model.
Goals
goals- Replace token, FLOP, benchmark score and 'one turn' as units of intelligence with a typed event 𝔍_IPM = (task, quality, semantic work, physical computation, scaffolding, metadata).
- Measure the relation physical computation → effective semantic work → verifiable quality, and compare systems on a Pareto frontier instead of a single score.
- Put the five falsifiable propositions — token hypothesis, FLOPs sufficiency, binary burden, scaffolding separation, semantic intermediate utility — in front of experiments in the order A → D → B → C → E.
- Publish under the IPM Minimum Reporting Standard: boundary, energy type, hidden work, uncertainty and measurement grade on every number.
Open questions
open_questions- μI has no operational identification yet (Experiment C); the series itself says μI must pass a predictive/explanatory utility test or be revised or eliminated.
- The pilot's quality axis measured output-format compliance for two of three tasks; task semantic quality, output-contract compliance, system completion reliability and verifier reliability have to be separated before scaling.
- Same-model verifier serialization failed in 3 of 18 verifier trials; candidates abandoned on abort are not yet accounted; telemetry sampling costs about 24 % of trial wall time.
- CODE tasks have no measured quality until the pilot runs inside a disposable sandbox with code evaluation enabled.
- Experiments B, C, D and E are declared, not run.
Milestones
milestones- 2026-09-02: EML-IPM v0.1 theoretical series complete, 10/10 papers with a SHA-256 manifest; canonical index with unified notation and the v0.2 experimental entry point.
- 2026-09-02: Experiment A protocol (EML-IPM-XA-01 v0.1, READY FOR PILOT).
- 2026-09-03: instrument packages XA-02 (30-task pack), XA-03 (telemetry logger), XA-04 (A0→A5 runner), XA-05 (36-trial synthetic smoke gate), XA-06 (real-model pilot gate), XA-06L (local execution handoff).
- 2026-09-03: first real-model pilot on a local Qwythos-9B-v2 — 36 trials, sealed REAL_MODEL_PILOT_INCOMPLETE (33/36 complete).
- 2026-09-07: diagnostic analysis of the pilot with seven instrument revisions required before XA-07.
Research lines and systems in this program
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
RES-2026-0101 Execution and physical line — what one answer costs in turns, semantic work, energy and computational spacetime | belongs_to | PRG-2026-0101 Intelligence Physical Metrology (IPM) | ACTIVE | REL-2026-0353 |
RES-2026-0102 Quality line — measuring outcome quality without asking humans for a score | belongs_to | PRG-2026-0101 Intelligence Physical Metrology (IPM) | ACTIVE | REL-2026-0354 |
RES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence event | belongs_to | PRG-2026-0101 Intelligence Physical Metrology (IPM) | ACTIVE | REL-2026-0355 |
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
PRG-2026-0101 Intelligence Physical Metrology (IPM) | released_as | ART-2026-0101 EML-IPM v0.1 canonical series package — 10 papers + canonical index + SHA-256 manifest (UTF-8 Markdown) artifact://evemisslab/intelligence-physical-metrology/IPM_v0.1_Canonical_Series_Package.zip | ACTIVE | REL-2026-0390 |
RES-2026-0101 Execution and physical line — what one answer costs in turns, semantic work, energy and computational spacetime | belongs_to | PRG-2026-0101 Intelligence Physical Metrology (IPM) | ACTIVE | REL-2026-0353 |
RES-2026-0102 Quality line — measuring outcome quality without asking humans for a score | belongs_to | PRG-2026-0101 Intelligence Physical Metrology (IPM) | ACTIVE | REL-2026-0354 |
RES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence event | belongs_to | PRG-2026-0101 Intelligence Physical Metrology (IPM) | ACTIVE | REL-2026-0355 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/programs/PRG-2026-0101/
- Machine-readable
/ai/programs/PRG-2026-0101/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports