EVEMISSLAB

BenchmarkBEN-2026-0002v0.1

AER-0 architecture-comparison suite (R1–R6)

Deterministic, executable comparisons that grow round by round: R1 — two facts (one stable, one changing at hours 8 and 16), eight queries over 17 simulated hours, baselines B1 stateless / B2 fixed 6-hour TTL / B3 memory + tools / B4 evented compound agent; R3 — scattered policy vs centralized application gate vs AER-ECT on provenance, stale-write, audit and dependency binding; R4 — policy-topology scaling over N ∈ {1, 4, 16, 64} callers × 6 rules; R5 — ten bypass scenarios across five threat layers; R6 — coherent full-DB forgery, anchor mutation, prefix truncation with/without a trusted head, fail-closed anchor, orphan anchor, adapter conformance.

Research status
EXPERIMENTAL being tested experimentally
Evidence level
E2 Controlled experiment
Version
0.1
Updated
2026-09-08
Created
2026-09-08
Domain
Evaluation, AI Architecture, Agent Systems
Program
PRG-2026-0001 Adaptive Epistemic Systems
Authors
Neo.K (EveMissLab)
AI collaborators
Sol (GPT-5.6, OpenAI ChatGPT)

Purpose

purpose
Ask a narrower question each round: what remains distinct once the baseline is allowed to be as good as AER?

Metrics

metrics
  • stale answers
  • recomputations
  • stable/volatile refreshes
  • workflow reuse
  • provenance/version-conflict witnesses
  • policy sites, rule placements, blast radius, migration edits
  • PREVENTED / OPEN_DETECTED / OPEN_UNDETECTED per scenario
  • test counts

Evaluation protocol

evaluation_protocol
Baselines are strengthened deliberately (B4 in R1, centralized gate in R3) so ordinary mechanisms are not attributed to AER; every round states supported and not-measured claims separately.

What it does not measure

limitations
  • Not a general-intelligence benchmark; no performance, cost, security-certification or production claim; the LangGraph round is source-grounded, not executed.

Relations

SourceRelationTargetStatusID
BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)evaluatesTHY-2026-0006 AER-ECT: a mandatory epistemic commit transaction boundaryACTIVEREL-2026-0089
BEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)evaluatesTHY-2026-0002 Canonical symbolic state and candidate → verify → commit authorityACTIVEREL-2026-0090
EXP-2026-0001 AER-0 MVP v0.1 closure: are the invariants executable?uses_benchmarkBEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)ACTIVEREL-2026-0100
EXP-2026-0002 R1 — deterministic semantics comparison against four baselinesuses_benchmarkBEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)ACTIVEREL-2026-0106
EXP-2026-0003 R2 — source-grounded structural comparison with LangGraph 1.2.11uses_benchmarkBEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)ACTIVEREL-2026-0117
EXP-2026-0004 R3 — epistemic commit transaction vs scattered and centralized application gatesuses_benchmarkBEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)ACTIVEREL-2026-0128
EXP-2026-0005 R4 — policy mutation surface: scattered governance vs one mandatory boundaryuses_benchmarkBEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)ACTIVEREL-2026-0138
EXP-2026-0006 R5 — complete mediation and bypass resistanceuses_benchmarkBEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)ACTIVEREL-2026-0143
EXP-2026-0007 R6 — external trust anchor, process-separated writer, adapter conformanceuses_benchmarkBEN-2026-0002 AER-0 architecture-comparison suite (R1–R6)ACTIVEREL-2026-0150

History and provenance

Canonical URL
https://evemisslab.com/ai/benchmarks/BEN-2026-0002/
Machine-readable
/ai/benchmarks/BEN-2026-0002/index.json
Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45
Provenance
source
EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at
2026-09-11
generator
tools/extract_aes/extract.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports