BenchmarkBEN-2026-0101v0.1
XA-02 — 30-task pilot pack for the A0→A5 scaffolding response
30 tasks — 10 math, 10 code, 10 constraint — with a public task file (prompt, output contract, pre-registered quality projection), a private reference file that must never enter model context, a deterministic evaluator and a 30/30 self-test. Math and constraint answers are one JSON object; code answers are Python source scored by hidden tests, and code execution is refused unless explicitly enabled inside an external sandbox. The pack's own words: not a general intelligence benchmark but a controlled instrument for measuring scaffolding response under Experiment A.
Purpose
purpose- Fixed task set with objective, pre-registered quality projections so that the same model can be run from native single pass (A0) to full agentic scaffold (A5) and the quality change attributed to scaffolding rather than to task drift.
Tasks
tasks- MATH-001 (math, easy)
- MATH-002 (math, easy)
- MATH-003 (math, easy)
- MATH-004 (math, medium)
- MATH-005 (math, medium)
- MATH-006 (math, medium)
- MATH-007 (math, medium)
- MATH-008 (math, hard)
- MATH-009 (math, hard)
- MATH-010 (math, hard)
- CODE-001 (code, easy)
- CODE-002 (code, easy)
- CODE-003 (code, easy)
- CODE-004 (code, medium)
- CODE-005 (code, medium)
- CODE-006 (code, medium)
- CODE-007 (code, medium)
- CODE-008 (code, hard)
- CODE-009 (code, hard)
- CODE-010 (code, hard)
- CON-001 (constraint, easy)
- CON-002 (constraint, easy)
- CON-003 (constraint, easy)
- CON-004 (constraint, medium)
- CON-005 (constraint, medium)
- CON-006 (constraint, medium)
- CON-007 (constraint, medium)
- CON-008 (constraint, hard)
- CON-009 (constraint, hard)
- CON-010 (constraint, hard)
Metrics
metricsmath and constraint- weighted exact fields on one JSON object; constraint tasks satisfied / m with fatal constraints as hard gate
code- hidden tests passed / tests (IPM_ALLOW_CODE_EXEC=1 required)
output contract- Return only one JSON object. Do not use Markdown fences.
Evaluation protocol
evaluation_protocol- evaluate.py is deterministic; reference_private.json is evaluator-private; validation_report.json records the 30/30 package self-test; manifest.json carries per-file SHA-256.
Baselines
baseline_resultsself-test- SELF-TESTED_30_OF_30
What it does not measure
limitations- Only three of the thirty tasks (MATH-003, CODE-001, CON-003) have been used, in the 36-trial gate matrices.
- The output contract makes 'quality' depend on format obedience: the first real pilot showed a correct answer scored 0 for not being JSON.
Recorded fields
tags- EML-IPM-XA-02
- v0.1
- SELF-TESTED_30_OF_30
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | evaluates | THY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladder | ACTIVE | REL-2026-0446 |
BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | released_as | ART-2026-0113 EML-IPM-XA-02 v0.1 — 30-task pilot pack and reference evaluators (private references separated) artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_XA02_30Task_PilotPack_v0.1.zip | ACTIVE | REL-2026-0447 |
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1) | uses_benchmark | BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | ACTIVE | REL-2026-0461 |
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic) | uses_benchmark | BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | ACTIVE | REL-2026-0467 |
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03) | uses_benchmark | BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding response | ACTIVE | REL-2026-0473 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/benchmarks/BEN-2026-0101/
- Machine-readable
/ai/benchmarks/BEN-2026-0101/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports