EVEMISSLAB
English

基準BEN-2026-0003v0.1

PACC-LLM 混合 A/B/C benchmark

三種條件跑在完全相同的候選池上:A 純生成器、B 硬驗證器、C PACC runtime(另有事後的 D「elastic」)。v0.1:7 類 × 28 題 × 8 池 × 112 候選(1,568 個共享池實例)。指標:明示硬約束遵守、衍生一致性、軟意圖滿足、長程保持、原始與有效新穎度、pattern entropy。v0.2:16 個手寫自然語言任務、相同計費預算、只有對條件盲的評審看得到 gold rubric。

研究狀態
EXPERIMENTAL 正在進行實驗驗證
證據等級
E2 受控實驗
版本
0.1
更新
2026-09-09
建立
2026-09-09
領域
Evaluation, Reasoning
計畫
PRG-2026-0001 自適應世界狀態系統的第一原理框架
作者
Neo.K (EveMissLab)
AI 協作
Sol (GPT-5.6, OpenAI ChatGPT)

目的

purpose
Separate 'more rejection' from genuine gains in derived dependency coherence, intent handling and constraint-satisfying novelty.

指標

metrics
  • explicit hard adherence
  • derived coherence
  • soft-intent satisfaction
  • long-horizon retention
  • raw novelty
  • valid novelty
  • pattern entropy

它沒有測什麼

limitations
  • A single-model judge is not human evaluation; nothing about real models is measured until a live run exists.

關係

來源關係目標狀態ID
BEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkevaluatesTHY-2026-0005 PACC 猜想——四層收斂階梯ACTIVEREL-2026-0091
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證uses_benchmarkBEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkACTIVEREL-2026-0314
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)uses_benchmarkBEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkACTIVEREL-2026-0321
EXP-2026-0023 PACC-Hybrid v0.2——第一次真實模型執行,本地 9B 開放權重模型uses_benchmarkBEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkACTIVEREL-2026-0327
EXP-2026-0024 PACC-Hybrid v0.2——真實模型第二次執行:四次重複與標籤無關的創造廣度uses_benchmarkBEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkACTIVEREL-2026-0335
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算uses_benchmarkBEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkACTIVEREL-2026-0344

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/benchmarks/BEN-2026-0003/
機器可讀
/ai/benchmarks/BEN-2026-0003/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at
2026-09-11
generator
tools/extract_aes/extract.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports