EVEMISSLAB
English

基準BEN-2026-0002v0.1

AER-0 架構比較套件(R1–R6)

逐輪成長的決定性可執行比較:R1——兩個事實(一穩定、一在第 8 與 16 小時改變)、17 個模擬小時內 8 次查詢,基線 B1 無狀態/B2 固定 6 小時 TTL/B3 記憶 + 工具/B4 事件驅動複合 agent;R3——分散政策 vs 集中式應用 gate vs AER-ECT,比 provenance、stale write、audit 與依賴綁定;R4——N ∈ {1, 4, 16, 64} 個 caller × 6 條規則的政策拓撲縮放;R5——五個威脅層上的十種繞過情境;R6——連貫的整庫偽造、anchor 竄改、有/無可信 head 的前綴截斷、fail-closed anchor、孤兒 anchor、adapter 一致性。

研究狀態
EXPERIMENTAL 正在進行實驗驗證
證據等級
E2 受控實驗
版本
0.1
更新
2026-09-08
建立
2026-09-08
領域
Evaluation, AI Architecture, Agent Systems
計畫
PRG-2026-0001 自適應世界狀態系統的第一原理框架
作者
Neo.K (EveMissLab)
AI 協作
Sol (GPT-5.6, OpenAI ChatGPT)

目的

purpose
Ask a narrower question each round: what remains distinct once the baseline is allowed to be as good as AER?

指標

metrics
  • stale answers
  • recomputations
  • stable/volatile refreshes
  • workflow reuse
  • provenance/version-conflict witnesses
  • policy sites, rule placements, blast radius, migration edits
  • PREVENTED / OPEN_DETECTED / OPEN_UNDETECTED per scenario
  • test counts

評估協定

evaluation_protocol
Baselines are strengthened deliberately (B4 in R1, centralized gate in R3) so ordinary mechanisms are not attributed to AER; every round states supported and not-measured claims separately.

它沒有測什麼

limitations
  • Not a general-intelligence benchmark; no performance, cost, security-certification or production claim; the LangGraph round is source-grounded, not executed.

關係

來源關係目標狀態ID
BEN-2026-0002 AER-0 架構比較套件(R1–R6)evaluatesTHY-2026-0006 AER-ECT:強制的認識論 commit 交易邊界ACTIVEREL-2026-0089
BEN-2026-0002 AER-0 架構比較套件(R1–R6)evaluatesTHY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限ACTIVEREL-2026-0090
EXP-2026-0001 AER-0 MVP v0.1 收束:不變量能不能被執行?uses_benchmarkBEN-2026-0002 AER-0 架構比較套件(R1–R6)ACTIVEREL-2026-0100
EXP-2026-0002 R1——對四個基線的決定性語義比較uses_benchmarkBEN-2026-0002 AER-0 架構比較套件(R1–R6)ACTIVEREL-2026-0106
EXP-2026-0003 R2——對 LangGraph 1.2.11 的 source-grounded 結構比較uses_benchmarkBEN-2026-0002 AER-0 架構比較套件(R1–R6)ACTIVEREL-2026-0117
EXP-2026-0004 R3——認識論 commit 交易 vs 分散式與集中式應用 gateuses_benchmarkBEN-2026-0002 AER-0 架構比較套件(R1–R6)ACTIVEREL-2026-0128
EXP-2026-0005 R4——政策變異面:分散治理 vs 單一強制邊界uses_benchmarkBEN-2026-0002 AER-0 架構比較套件(R1–R6)ACTIVEREL-2026-0138
EXP-2026-0006 R5——完全中介與繞過抗性uses_benchmarkBEN-2026-0002 AER-0 架構比較套件(R1–R6)ACTIVEREL-2026-0143
EXP-2026-0007 R6——外部信任 anchor、程序分離 writer、adapter 一致性uses_benchmarkBEN-2026-0002 AER-0 架構比較套件(R1–R6)ACTIVEREL-2026-0150

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/benchmarks/BEN-2026-0002/
機器可讀
/ai/benchmarks/BEN-2026-0002/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at
2026-09-11
generator
tools/extract_aes/extract.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports