基準BEN-2026-0101v0.1
XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包
30 題——數學 10、程式 10、約束 10——含公開任務檔(題目、輸出契約、預先登記的品質投影)、絕不可進入模型上下文的私有參考檔、確定性評分器與 30/30 自測。數學與約束題回一個 JSON 物件;程式題回 Python 原始碼、以隱藏測試評分,且除非在外部沙箱明確啟用,否則拒絕執行程式。套件自己的說法:不是通用智能 benchmark,而是 Experiment A 下量測鷹架響應的受控儀器。
目的
purpose- Fixed task set with objective, pre-registered quality projections so that the same model can be run from native single pass (A0) to full agentic scaffold (A5) and the quality change attributed to scaffolding rather than to task drift.
任務
tasks- MATH-001 (math, easy)
- MATH-002 (math, easy)
- MATH-003 (math, easy)
- MATH-004 (math, medium)
- MATH-005 (math, medium)
- MATH-006 (math, medium)
- MATH-007 (math, medium)
- MATH-008 (math, hard)
- MATH-009 (math, hard)
- MATH-010 (math, hard)
- CODE-001 (code, easy)
- CODE-002 (code, easy)
- CODE-003 (code, easy)
- CODE-004 (code, medium)
- CODE-005 (code, medium)
- CODE-006 (code, medium)
- CODE-007 (code, medium)
- CODE-008 (code, hard)
- CODE-009 (code, hard)
- CODE-010 (code, hard)
- CON-001 (constraint, easy)
- CON-002 (constraint, easy)
- CON-003 (constraint, easy)
- CON-004 (constraint, medium)
- CON-005 (constraint, medium)
- CON-006 (constraint, medium)
- CON-007 (constraint, medium)
- CON-008 (constraint, hard)
- CON-009 (constraint, hard)
- CON-010 (constraint, hard)
指標
metricsmath and constraint- weighted exact fields on one JSON object; constraint tasks satisfied / m with fatal constraints as hard gate
code- hidden tests passed / tests (IPM_ALLOW_CODE_EXEC=1 required)
output contract- Return only one JSON object. Do not use Markdown fences.
評估協定
evaluation_protocol- evaluate.py is deterministic; reference_private.json is evaluator-private; validation_report.json records the 30/30 package self-test; manifest.json carries per-file SHA-256.
基線
baseline_resultsself-test- SELF-TESTED_30_OF_30
它沒有測什麼
limitations- Only three of the thirty tasks (MATH-003, CODE-001, CON-003) have been used, in the 36-trial gate matrices.
- The output contract makes 'quality' depend on format obedience: the first real pilot showed a correct answer scored 0 for not being JSON.
記錄欄位
tags- EML-IPM-XA-02
- v0.1
- SELF-TESTED_30_OF_30
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包 | evaluates | THY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯 | ACTIVE | REL-2026-0446 |
BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包 | released_as | ART-2026-0113 EML-IPM-XA-02 v0.1 — 30-task pilot pack and reference evaluators (private references separated) artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_XA02_30Task_PilotPack_v0.1.zip | ACTIVE | REL-2026-0447 |
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1 | uses_benchmark | BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包 | ACTIVE | REL-2026-0461 |
EXP-2026-0102 XA-05——以腳本化供應商跑的 36 試驗端到端煙霧閘(合成) | uses_benchmark | BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包 | ACTIVE | REL-2026-0467 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | uses_benchmark | BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包 | ACTIVE | REL-2026-0473 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/benchmarks/BEN-2026-0101/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports