AI Research Laboratory4 records
Benchmarks
Each benchmark states what it measures and what it does not.
- BenchmarkBEN-2026-0001PACC micro-lab protocol v0.1 (frozen gates)The preregistered decision rule reused unchanged from v0.1 to v0.13: E0 purity, held-out representation distance D_R = E[JS(Φ(S_N), S_P)] ≤ 0.05, update commutation D_U ≤ 0.05, off-manifold intervention D_I ≤ 0.08, action agreement ≥ 0.85, a shuffled-target negative control the real map must beat, a design-independence rule (agreement ≥ 0.999 and R² ≥ 0.995 collapse two designs into one family), verdict levels 0–3, and a strong PACC-A gate of at least three independent convergent families.
- BenchmarkBEN-2026-0003PACC-LLM Hybrid A/B/C benchmarkThree conditions over identical candidate pools: A generator-only, B hard verifier, C PACC runtime (plus post-hoc D 'elastic'). v0.1: 7 categories × 28 tasks × 8 pools × 112 candidates (1,568 shared pool instances). Metrics: explicit hard adherence, derived coherence, soft-intent satisfaction, long-horizon retention, raw and valid novelty, pattern entropy. v0.2: 16 hand-authored natural-language tasks, equal accounted budgets, gold rubric visible only to a condition-blind judge.
- BenchmarkBEN-2026-0002AER-0 architecture-comparison suite (R1–R6)Deterministic, executable comparisons that grow round by round: R1 — two facts (one stable, one changing at hours 8 and 16), eight queries over 17 simulated hours, baselines B1 stateless / B2 fixed 6-hour TTL / B3 memory + tools / B4 evented compound agent; R3 — scattered policy vs centralized application gate vs AER-ECT on provenance, stale-write, audit and dependency binding; R4 — policy-topology scaling over N ∈ {1, 4, 16, 64} callers × 6 rules; R5 — ten bypass scenarios across five threat layers; R6 — coherent full-DB forgery, anchor mutation, prefix truncation with/without a trusted head, fail-closed anchor, orphan anchor, adapter conformance.
- BenchmarkBEN-2026-0101XA-02 — 30-task pilot pack for the A0→A5 scaffolding response30 tasks — 10 math, 10 code, 10 constraint — with a public task file (prompt, output contract, pre-registered quality projection), a private reference file that must never enter model context, a deterministic evaluator and a 30/30 self-test. Math and constraint answers are one JSON object; code answers are Python source scored by hidden tests, and code execution is refused unless explicitly enabled inside an external sandbox. The pack's own words: not a general intelligence benchmark but a controlled instrument for measuring scaffolding response under Experiment A.