EVEMISSLAB

ResearchRES-2026-0103v0.1

Capability line — how much intelligence remains without the loop, and the unified intelligence event

Papers 09–10 and the v0.2 Experiment A instrument. A system is (model, scaffolding vector); single-pass quality Q_SP and full-system quality Q_F give the scaffolding survival ratio SSR = Q_SP/Q_F, the dependence ratio SDR, the scaffold cost multiplier SCM and marginal scaffolding yields along a controlled ablation ladder A0…A5 — all read together with the physical overhead, because a loop is not cheating; hiding its cost is. Paper 10 packs quality, semantic work, physical computation and scaffolding into the canonical intelligence event, compares systems on Pareto frontiers under a no-premature-scalarization rule, and fixes a minimum reporting standard. The line's first real data point is the 2026-09-03 pilot on a local 9B model.

Research status
EXPERIMENTAL being tested experimentally
Evidence level
E2 Controlled experiment
Version
0.1
Updated
2026-09-07
Created
2026-09-02
Domain
Evaluation, Agent Systems, Computation
Program
PRG-2026-0101 Intelligence Physical Metrology (IPM)
Authors
Neo.K (EveMissLab)
AI collaborators
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT)

Research questions

research_questions
  • Of a final answer, how much came from the model's native single pass and how much from retries, sampling, verifiers, tools, memory and planners — at what physical cost?
  • Does a stable scaffolding response curve exist for a fixed model and task set, and where does it enter the brute-force region?
  • Can two systems with the same final quality be told apart by capability source and physical cost rather than by a leaderboard score?

Claims

claims
  • System capability ≠ model-native capability; Pass@k ≠ Pass@1; tool access ≠ tool utilization intelligence; invisible output ≠ zero cost.
  • SSR/SDR describe the structure of an intelligence source, not a defect; they must be read with SCM and the scaffolding physical overhead.
  • Intelligence is not any single element of the tuple (task, quality, semantic work, physical computation, scaffolding, metadata); the research object is the relation P_compute → N_μ → 𝔔.
  • First real pilot (three easy tasks, local 9B, 2026-09-03): SSR = 1.0 with a 3.1× device-energy multiplier from A0 to A5 — and the instrument's quality axis turned out to measure output-format compliance on two of the three tasks.

Limitations

limitations
  • One model, three tasks, two replicates, same-model verifier, code quality unmeasured, gate INCOMPLETE — a dataset for instrument revision, not an estimate.
  • Shapley-style scaffold attribution, information-matched tool controls and the 900-trial physical comparison are all still ahead.

Relations

SourceRelationTargetStatusID
RES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventbelongs_toPRG-2026-0101 Intelligence Physical Metrology (IPM)ACTIVEREL-2026-0355
RES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventextendsRES-2026-0101 Execution and physical line — what one answer costs in turns, semantic work, energy and computational spacetimeACTIVEREL-2026-0356
RES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventextendsRES-2026-0102 Quality line — measuring outcome quality without asking humans for a scoreACTIVEREL-2026-0357
RES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventdevelopsTHY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladderACTIVEREL-2026-0366
RES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventdevelopsTHY-2026-0110 The canonical intelligence event, Pareto comparison and no premature scalarizationACTIVEREL-2026-0367
CLM-2026-0104 F4 — Scaffolding separationbelongs_toRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0386
CLM-2026-0105 F5 — Semantic intermediate utilitybelongs_toRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0388
PAP-2026-0109 Paper 09 — How much intelligence remains without the loop? Single-pass capability, scaffolding dependence, and hidden computational costreportsRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0423
PAP-2026-0110 Paper 10 — How much physical world does an answer cost? A unified metrology framework for intelligence yieldreportsRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0427
PAP-2026-0111 IPM v0.1 canonical index — series overview, unified notation and the v0.2 experimental entry pointreportsRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0433
SYS-2026-0102 XA-04 — A0→A5 model runner and scaffold orchestratorsupportsRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0453
SYS-2026-0103 XA-06 — real-model pilot gatesupportsRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0456
SYS-2026-0104 XA-06L — local real-model execution handoff packsupportsRES-2026-0103 Capability line — how much intelligence remains without the loop, and the unified intelligence eventACTIVEREL-2026-0459

History and provenance

Canonical URL
https://evemisslab.com/ai/research/RES-2026-0103/
Machine-readable
/ai/research/RES-2026-0103/index.json
Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45
Provenance
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports