AI Research Laboratory5 records
Claims
Claims as first-class objects, with what supports and what contradicts them.
- ClaimCLM-2026-0101F1 — Token hypothesisIf the token is a good universal unit of intelligent work, then N_μ^eff / TokenCount should be approximately stable across models, languages and phrasings once task quality is controlled. Large drift across models, modalities or expression forms weakens 'token = intelligent work unit' to an implementation/interface proxy.
- ClaimCLM-2026-0102F2 — FLOPs sufficiencyIf FLOPs suffice to describe physical computational cost, then with operations controlled, wall time, energy, memory traffic, interconnect traffic and memory residency should not show large independent variation. If same-FLOPs workloads differ greatly because of memory patterns or topology, a single FLOPs cost model is insufficient.
- ClaimCLM-2026-0103F3 — Binary burden hypothesisIf direct numeric human rating were already the minimum-burden, high-quality measurement interface, then well-designed binary / pairwise adaptive protocols should have no advantage in response time, consistency, dropout, predictive validity or fatigue. BRQM predicts an advantage in at least some settings; it can be tested by randomizing participants across direct 0–10, structured yes/no and adaptive pairwise formats.
- ClaimCLM-2026-0104F4 — Scaffolding separationIf scaffolding does not change the structure of the capability source, SSR ≈ 1 should hold across most tasks and test-time compute budgets. If a large share of benchmark quality appears only under multi-sample, tool, verifier or loop conditions, the distinction between model-native and system capability is empirically necessary. First data point: three easy tasks on a local 9B model gave SSR = 1.0 — consistent with 'no gap' on tasks the native pass already solves, and uninformative beyond that because the quality axis was confounded by format compliance.
- ClaimCLM-2026-0105F5 — Semantic intermediate utilityIf the semantic middle layer N_μ has no measurement value, then adding it should not improve the explanation or prediction of cross-architecture efficiency, error paths, scaffold gain or task transfer over a direct physical-cost → quality model. μI itself must pass this predictive/explanatory utility test or be revised or eliminated.