BenchmarkBEN-2026-0003v0.1
PACC-LLM Hybrid A/B/C benchmark
Three conditions over identical candidate pools: A generator-only, B hard verifier, C PACC runtime (plus post-hoc D 'elastic'). v0.1: 7 categories × 28 tasks × 8 pools × 112 candidates (1,568 shared pool instances). Metrics: explicit hard adherence, derived coherence, soft-intent satisfaction, long-horizon retention, raw and valid novelty, pattern entropy. v0.2: 16 hand-authored natural-language tasks, equal accounted budgets, gold rubric visible only to a condition-blind judge.
Purpose
purpose- Separate 'more rejection' from genuine gains in derived dependency coherence, intent handling and constraint-satisfying novelty.
Metrics
metrics- explicit hard adherence
- derived coherence
- soft-intent satisfaction
- long-horizon retention
- raw novelty
- valid novelty
- pattern entropy
What it does not measure
limitations- A single-model judge is not human evaluation; nothing about real models is measured until a live run exists.
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
BEN-2026-0003 PACC-LLM Hybrid A/B/C benchmark | evaluates | THY-2026-0005 PACC conjecture — the four-level convergence ladder | ACTIVE | REL-2026-0091 |
EXP-2026-0021 PACC-Hybrid v0.1 — synthetic A/B/C architecture witness | uses_benchmark | BEN-2026-0003 PACC-LLM Hybrid A/B/C benchmark | ACTIVE | REL-2026-0314 |
EXP-2026-0022 PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed) | uses_benchmark | BEN-2026-0003 PACC-LLM Hybrid A/B/C benchmark | ACTIVE | REL-2026-0321 |
EXP-2026-0023 PACC-Hybrid v0.2 — first real-model run, on a local 9B open-weight model | uses_benchmark | BEN-2026-0003 PACC-LLM Hybrid A/B/C benchmark | ACTIVE | REL-2026-0327 |
EXP-2026-0024 PACC-Hybrid v0.2 — real-model run 2: four repetitions and label-free creative breadth | uses_benchmark | BEN-2026-0003 PACC-LLM Hybrid A/B/C benchmark | ACTIVE | REL-2026-0335 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | uses_benchmark | BEN-2026-0003 PACC-LLM Hybrid A/B/C benchmark | ACTIVE | REL-2026-0344 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/benchmarks/BEN-2026-0003/
- Machine-readable
/ai/benchmarks/BEN-2026-0003/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports