EVEMISSLAB

BenchmarkBEN-2026-0101v0.1

XA-02 — 30-task pilot pack for the A0→A5 scaffolding response

30 tasks — 10 math, 10 code, 10 constraint — with a public task file (prompt, output contract, pre-registered quality projection), a private reference file that must never enter model context, a deterministic evaluator and a 30/30 self-test. Math and constraint answers are one JSON object; code answers are Python source scored by hidden tests, and code execution is refused unless explicitly enabled inside an external sandbox. The pack's own words: not a general intelligence benchmark but a controlled instrument for measuring scaffolding response under Experiment A.

Research status
STABLE the current conclusions are relatively stable
Evidence level
E2 Controlled experiment
Version
0.1
Updated
2026-09-02
Created
2026-09-02
Domain
Evaluation
Program
PRG-2026-0101 Intelligence Physical Metrology (IPM)
Authors
Neo.K (EveMissLab)
AI collaborators
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT)

Purpose

purpose
Fixed task set with objective, pre-registered quality projections so that the same model can be run from native single pass (A0) to full agentic scaffold (A5) and the quality change attributed to scaffolding rather than to task drift.

Tasks

tasks
  • MATH-001 (math, easy)
  • MATH-002 (math, easy)
  • MATH-003 (math, easy)
  • MATH-004 (math, medium)
  • MATH-005 (math, medium)
  • MATH-006 (math, medium)
  • MATH-007 (math, medium)
  • MATH-008 (math, hard)
  • MATH-009 (math, hard)
  • MATH-010 (math, hard)
  • CODE-001 (code, easy)
  • CODE-002 (code, easy)
  • CODE-003 (code, easy)
  • CODE-004 (code, medium)
  • CODE-005 (code, medium)
  • CODE-006 (code, medium)
  • CODE-007 (code, medium)
  • CODE-008 (code, hard)
  • CODE-009 (code, hard)
  • CODE-010 (code, hard)
  • CON-001 (constraint, easy)
  • CON-002 (constraint, easy)
  • CON-003 (constraint, easy)
  • CON-004 (constraint, medium)
  • CON-005 (constraint, medium)
  • CON-006 (constraint, medium)
  • CON-007 (constraint, medium)
  • CON-008 (constraint, hard)
  • CON-009 (constraint, hard)
  • CON-010 (constraint, hard)

Metrics

metrics
math and constraint
weighted exact fields on one JSON object; constraint tasks satisfied / m with fatal constraints as hard gate
code
hidden tests passed / tests (IPM_ALLOW_CODE_EXEC=1 required)
output contract
Return only one JSON object. Do not use Markdown fences.

Evaluation protocol

evaluation_protocol
evaluate.py is deterministic; reference_private.json is evaluator-private; validation_report.json records the 30/30 package self-test; manifest.json carries per-file SHA-256.

Baselines

baseline_results
self-test
SELF-TESTED_30_OF_30

What it does not measure

limitations
  • Only three of the thirty tasks (MATH-003, CODE-001, CON-003) have been used, in the 36-trial gate matrices.
  • The output contract makes 'quality' depend on format obedience: the first real pilot showed a correct answer scored 0 for not being JSON.

Recorded fields

tags
  • EML-IPM-XA-02
  • v0.1
  • SELF-TESTED_30_OF_30

Relations

SourceRelationTargetStatusID
BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding responseevaluatesTHY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladderACTIVEREL-2026-0446
BEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding responsereleased_asART-2026-0113 EML-IPM-XA-02 v0.1 — 30-task pilot pack and reference evaluators (private references separated) artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_XA02_30Task_PilotPack_v0.1.zipACTIVEREL-2026-0447
EXP-2026-0101 Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)uses_benchmarkBEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding responseACTIVEREL-2026-0461
EXP-2026-0102 XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic)uses_benchmarkBEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding responseACTIVEREL-2026-0467
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03)uses_benchmarkBEN-2026-0101 XA-02 — 30-task pilot pack for the A0→A5 scaffolding responseACTIVEREL-2026-0473

History and provenance

Canonical URL
https://evemisslab.com/ai/benchmarks/BEN-2026-0101/
Machine-readable
/ai/benchmarks/BEN-2026-0101/index.json
Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45
Provenance
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports