基準BEN-2026-0003v0.1
PACC-LLM 混合 A/B/C benchmark
三種條件跑在完全相同的候選池上:A 純生成器、B 硬驗證器、C PACC runtime(另有事後的 D「elastic」)。v0.1:7 類 × 28 題 × 8 池 × 112 候選(1,568 個共享池實例)。指標:明示硬約束遵守、衍生一致性、軟意圖滿足、長程保持、原始與有效新穎度、pattern entropy。v0.2:16 個手寫自然語言任務、相同計費預算、只有對條件盲的評審看得到 gold rubric。
目的
purpose- Separate 'more rejection' from genuine gains in derived dependency coherence, intent handling and constraint-satisfying novelty.
指標
metrics- explicit hard adherence
- derived coherence
- soft-intent satisfaction
- long-horizon retention
- raw novelty
- valid novelty
- pattern entropy
它沒有測什麼
limitations- A single-model judge is not human evaluation; nothing about real models is measured until a live run exists.
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | evaluates | THY-2026-0005 PACC 猜想——四層收斂階梯 | ACTIVE | REL-2026-0091 |
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證 | uses_benchmark | BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | ACTIVE | REL-2026-0314 |
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行) | uses_benchmark | BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | ACTIVE | REL-2026-0321 |
EXP-2026-0023 PACC-Hybrid v0.2——第一次真實模型執行,本地 9B 開放權重模型 | uses_benchmark | BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | ACTIVE | REL-2026-0327 |
EXP-2026-0024 PACC-Hybrid v0.2——真實模型第二次執行:四次重複與標籤無關的創造廣度 | uses_benchmark | BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | ACTIVE | REL-2026-0335 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | uses_benchmark | BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | ACTIVE | REL-2026-0344 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/benchmarks/BEN-2026-0003/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports