基準BEN-2026-0002v0.1
AER-0 架構比較套件(R1–R6)
逐輪成長的決定性可執行比較:R1——兩個事實(一穩定、一在第 8 與 16 小時改變)、17 個模擬小時內 8 次查詢,基線 B1 無狀態/B2 固定 6 小時 TTL/B3 記憶 + 工具/B4 事件驅動複合 agent;R3——分散政策 vs 集中式應用 gate vs AER-ECT,比 provenance、stale write、audit 與依賴綁定;R4——N ∈ {1, 4, 16, 64} 個 caller × 6 條規則的政策拓撲縮放;R5——五個威脅層上的十種繞過情境;R6——連貫的整庫偽造、anchor 竄改、有/無可信 head 的前綴截斷、fail-closed anchor、孤兒 anchor、adapter 一致性。
目的
purpose- Ask a narrower question each round: what remains distinct once the baseline is allowed to be as good as AER?
指標
metrics- stale answers
- recomputations
- stable/volatile refreshes
- workflow reuse
- provenance/version-conflict witnesses
- policy sites, rule placements, blast radius, migration edits
- PREVENTED / OPEN_DETECTED / OPEN_UNDETECTED per scenario
- test counts
評估協定
evaluation_protocol- Baselines are strengthened deliberately (B4 in R1, centralized gate in R3) so ordinary mechanisms are not attributed to AER; every round states supported and not-measured claims separately.
它沒有測什麼
limitations- Not a general-intelligence benchmark; no performance, cost, security-certification or production claim; the LangGraph round is source-grounded, not executed.
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
BEN-2026-0002 AER-0 架構比較套件(R1–R6) | evaluates | THY-2026-0006 AER-ECT:強制的認識論 commit 交易邊界 | ACTIVE | REL-2026-0089 |
BEN-2026-0002 AER-0 架構比較套件(R1–R6) | evaluates | THY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限 | ACTIVE | REL-2026-0090 |
EXP-2026-0001 AER-0 MVP v0.1 收束:不變量能不能被執行? | uses_benchmark | BEN-2026-0002 AER-0 架構比較套件(R1–R6) | ACTIVE | REL-2026-0100 |
EXP-2026-0002 R1——對四個基線的決定性語義比較 | uses_benchmark | BEN-2026-0002 AER-0 架構比較套件(R1–R6) | ACTIVE | REL-2026-0106 |
EXP-2026-0003 R2——對 LangGraph 1.2.11 的 source-grounded 結構比較 | uses_benchmark | BEN-2026-0002 AER-0 架構比較套件(R1–R6) | ACTIVE | REL-2026-0117 |
EXP-2026-0004 R3——認識論 commit 交易 vs 分散式與集中式應用 gate | uses_benchmark | BEN-2026-0002 AER-0 架構比較套件(R1–R6) | ACTIVE | REL-2026-0128 |
EXP-2026-0005 R4——政策變異面:分散治理 vs 單一強制邊界 | uses_benchmark | BEN-2026-0002 AER-0 架構比較套件(R1–R6) | ACTIVE | REL-2026-0138 |
EXP-2026-0006 R5——完全中介與繞過抗性 | uses_benchmark | BEN-2026-0002 AER-0 架構比較套件(R1–R6) | ACTIVE | REL-2026-0143 |
EXP-2026-0007 R6——外部信任 anchor、程序分離 writer、adapter 一致性 | uses_benchmark | BEN-2026-0002 AER-0 架構比較套件(R1–R6) | ACTIVE | REL-2026-0150 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/benchmarks/BEN-2026-0002/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports