EVEMISSLAB
English

基準BEN-2026-0101v0.1

XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包

30 題——數學 10、程式 10、約束 10——含公開任務檔(題目、輸出契約、預先登記的品質投影)、絕不可進入模型上下文的私有參考檔、確定性評分器與 30/30 自測。數學與約束題回一個 JSON 物件;程式題回 Python 原始碼、以隱藏測試評分,且除非在外部沙箱明確啟用,否則拒絕執行程式。套件自己的說法:不是通用智能 benchmark,而是 Experiment A 下量測鷹架響應的受控儀器。

研究狀態
STABLE 目前的研究結論相對穩定
證據等級
E2 受控實驗
版本
0.1
更新
2026-09-02
建立
2026-09-02
領域
Evaluation
計畫
PRG-2026-0101 智能的物理計量(IPM)
作者
Neo.K (EveMissLab)
AI 協作
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT)

目的

purpose
Fixed task set with objective, pre-registered quality projections so that the same model can be run from native single pass (A0) to full agentic scaffold (A5) and the quality change attributed to scaffolding rather than to task drift.

任務

tasks
  • MATH-001 (math, easy)
  • MATH-002 (math, easy)
  • MATH-003 (math, easy)
  • MATH-004 (math, medium)
  • MATH-005 (math, medium)
  • MATH-006 (math, medium)
  • MATH-007 (math, medium)
  • MATH-008 (math, hard)
  • MATH-009 (math, hard)
  • MATH-010 (math, hard)
  • CODE-001 (code, easy)
  • CODE-002 (code, easy)
  • CODE-003 (code, easy)
  • CODE-004 (code, medium)
  • CODE-005 (code, medium)
  • CODE-006 (code, medium)
  • CODE-007 (code, medium)
  • CODE-008 (code, hard)
  • CODE-009 (code, hard)
  • CODE-010 (code, hard)
  • CON-001 (constraint, easy)
  • CON-002 (constraint, easy)
  • CON-003 (constraint, easy)
  • CON-004 (constraint, medium)
  • CON-005 (constraint, medium)
  • CON-006 (constraint, medium)
  • CON-007 (constraint, medium)
  • CON-008 (constraint, hard)
  • CON-009 (constraint, hard)
  • CON-010 (constraint, hard)

指標

metrics
math and constraint
weighted exact fields on one JSON object; constraint tasks satisfied / m with fatal constraints as hard gate
code
hidden tests passed / tests (IPM_ALLOW_CODE_EXEC=1 required)
output contract
Return only one JSON object. Do not use Markdown fences.

評估協定

evaluation_protocol
evaluate.py is deterministic; reference_private.json is evaluator-private; validation_report.json records the 30/30 package self-test; manifest.json carries per-file SHA-256.

基線

baseline_results
self-test
SELF-TESTED_30_OF_30

它沒有測什麼

limitations
  • Only three of the thirty tasks (MATH-003, CODE-001, CON-003) have been used, in the 36-trial gate matrices.
  • The output contract makes 'quality' depend on format obedience: the first real pilot showed a correct answer scored 0 for not being JSON.

記錄欄位

tags
  • EML-IPM-XA-02
  • v0.1
  • SELF-TESTED_30_OF_30

關係

來源關係目標狀態ID
BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包evaluatesTHY-2026-0109 鷹架能力紀錄:SSR、SDR、SCM 與消融階梯ACTIVEREL-2026-0446
BEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包released_asART-2026-0113 EML-IPM-XA-02 v0.1 — 30-task pilot pack and reference evaluators (private references separated) artifact://evemisslab/intelligence-physical-metrology/IPM_v0.2_XA02_30Task_PilotPack_v0.1.zipACTIVEREL-2026-0447
EXP-2026-0101 Experiment A——單次智能與鷹架增益的受控實驗協定 v0.1uses_benchmarkBEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包ACTIVEREL-2026-0461
EXP-2026-0102 XA-05——以腳本化供應商跑的 36 試驗端到端煙霧閘(合成)uses_benchmarkBEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包ACTIVEREL-2026-0467
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03)uses_benchmarkBEN-2026-0101 XA-02——A0→A5 鷹架響應的 30 題 pilot 任務包ACTIVEREL-2026-0473

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/benchmarks/BEN-2026-0101/
機器可讀
/ai/benchmarks/BEN-2026-0101/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports