ClaimCLM-2026-0104v0.1
F4 — Scaffolding separation
If scaffolding does not change the structure of the capability source, SSR ≈ 1 should hold across most tasks and test-time compute budgets. If a large share of benchmark quality appears only under multi-sample, tool, verifier or loop conditions, the distinction between model-native and system capability is empirically necessary. First data point: three easy tasks on a local 9B model gave SSR = 1.0 — consistent with 'no gap' on tasks the native pass already solves, and uninformative beyond that because the quality axis was confounded by format compliance.
Recorded fields
falsification_conditions- Falsified for the separation's necessity if Q_F ≈ Q_SP holds across most tasks, models and budgets; supported if the gap is common. The 2026-09-03 pilot is one uninformative-to-weak point on the 'no gap' side.
tags- falsifiable proposition
- IPM v0.1 canonical index §23
- Paper 10
Relations
History and provenance
- Canonical URL
- https://evemisslab.com/ai/claims/CLM-2026-0104/
- Machine-readable
/ai/claims/CLM-2026-0104/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports