BenchmarkBEN-2026-0002v0.1
AER-0 architecture-comparison suite (R1–R6)
Deterministic, executable comparisons that grow round by round: R1 — two facts (one stable, one changing at hours 8 and 16), eight queries over 17 simulated hours, baselines B1 stateless / B2 fixed 6-hour TTL / B3 memory + tools / B4 evented compound agent; R3 — scattered policy vs centralized application gate vs AER-ECT on provenance, stale-write, audit and dependency binding; R4 — policy-topology scaling over N ∈ {1, 4, 16, 64} callers × 6 rules; R5 — ten bypass scenarios across five threat layers; R6 — coherent full-DB forgery, anchor mutation, prefix truncation with/without a trusted head, fail-closed anchor, orphan anchor, adapter conformance.
Purpose
purpose- Ask a narrower question each round: what remains distinct once the baseline is allowed to be as good as AER?
Metrics
metrics- stale answers
- recomputations
- stable/volatile refreshes
- workflow reuse
- provenance/version-conflict witnesses
- policy sites, rule placements, blast radius, migration edits
- PREVENTED / OPEN_DETECTED / OPEN_UNDETECTED per scenario
- test counts
Evaluation protocol
evaluation_protocol- Baselines are strengthened deliberately (B4 in R1, centralized gate in R3) so ordinary mechanisms are not attributed to AER; every round states supported and not-measured claims separately.
What it does not measure
limitations- Not a general-intelligence benchmark; no performance, cost, security-certification or production claim; the LangGraph round is source-grounded, not executed.
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | evaluates | THY-2026-0006 AER-ECT: a mandatory epistemic commit transaction boundary | ACTIVE | REL-2026-0089 |
BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | evaluates | THY-2026-0002 Canonical symbolic state and candidate → verify → commit authority | ACTIVE | REL-2026-0090 |
EXP-2026-0001 AER-0 MVP v0.1 closure: are the invariants executable? | uses_benchmark | BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | ACTIVE | REL-2026-0100 |
EXP-2026-0002 R1 — deterministic semantics comparison against four baselines | uses_benchmark | BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | ACTIVE | REL-2026-0106 |
EXP-2026-0003 R2 — source-grounded structural comparison with LangGraph 1.2.11 | uses_benchmark | BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | ACTIVE | REL-2026-0117 |
EXP-2026-0004 R3 — epistemic commit transaction vs scattered and centralized application gates | uses_benchmark | BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | ACTIVE | REL-2026-0128 |
EXP-2026-0005 R4 — policy mutation surface: scattered governance vs one mandatory boundary | uses_benchmark | BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | ACTIVE | REL-2026-0138 |
EXP-2026-0006 R5 — complete mediation and bypass resistance | uses_benchmark | BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | ACTIVE | REL-2026-0143 |
EXP-2026-0007 R6 — external trust anchor, process-separated writer, adapter conformance | uses_benchmark | BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6) | ACTIVE | REL-2026-0150 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/benchmarks/BEN-2026-0002/
- Machine-readable
/ai/benchmarks/BEN-2026-0002/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports