EVEMISSLAB
English

實驗EXP-2026-0021v0.1

PACC-Hybrid v0.1——合成 A/B/C 架構見證

在 1,568 個完全相同的候選池上,純生成器(A)、硬驗證器(B)、PACC runtime(C)以硬約束遵守、衍生一致性、軟意圖滿足、長程保持、新穎度與 pattern entropy 計分。B 已經把字面硬約束遵守飽和到 1.0;C 相對 B 衍生一致性 +0.2085、長程保持 +0.0142、有效新穎度 +0.1964、軟意圖滿足 −0.0076,並在純創意控制組中提高逐點新穎度(0.9864 vs 0.8889)卻讓 pattern entropy 塌縮(0.2284 vs 0.7199)。事後的「elastic」探索策略(D)在衍生一致性仍為 1.0 下把 entropy 恢復到 0.7638。零次真實 LLM 呼叫;八個 seed 符號穩定。

研究狀態
STABLE 目前的研究結論相對穩定
證據等級
E3 重複實驗
結果
MIXED
資料基礎
SYNTHETIC 合成數據與理論推理。現在很多人把合成數據當成真的;這個實驗室刻意反過來說——在真正的混合模型出現之前,推論就只是推論,理論上可能不等於實際上可能。
版本
0.1
更新
2026-09-09
建立
2026-09-09
領域
Reasoning, Evaluation
計畫
PRG-2026-0001 自適應世界狀態系統的第一原理框架
作者
Neo.K (EveMissLab)
AI 協作
Sol (GPT-5.6, OpenAI ChatGPT)

假設

hypothesis
A PACC-style commit space changes selection quality beyond what a hard verifier achieves, and any creative-breadth loss is separable from the commit constraints.

設定

dataset_ids
benchmark_ids
software_environment
Python; synthetic generators and scorers; no network, no LLM.

執行

run_count
9
random_seeds
  • 20260909
  • 8 fixed secondary seeds
controls
  • identical candidate pools for A/B/C
  • pure-creative negative control
  • post-hoc elastic diagnostic labelled as such
metrics
verdict
SYNTHETIC_PARETO_WITNESS_REASONING_UP_BREADTH_COLLAPSE_MITIGATABLE
scale
7 categories × 28 tasks × 8 pools × 112 candidates; 1568 pool instances; 0 LLM calls
overall
A
hard
0.9936
derived
0.7902
soft
0.7878
long_horizon
0.9648
valid_novelty
0.5487
B
hard
1.0
derived
0.7915
soft
0.7859
long_horizon
0.9676
valid_novelty
0.551
C
hard
1.0
derived
1.0
soft
0.7783
long_horizon
0.9818
valid_novelty
0.7474
pure_creative
A
raw_novelty
0.8889
pattern_entropy
0.7199
C
raw_novelty
0.9864
pattern_entropy
0.2284
D_elastic_posthoc
raw_novelty
0.9225
pattern_entropy
0.7638
eight_seed_deltas
C_minus_B_derived
mean
0.1865
sign
8/8 positive
C_minus_A_pattern_entropy
mean
-0.3894
sign
8/8 negative
B_literal_fallback_rate
0.286

詮釋

interpretation
Not a simple reasoning-up / imagination-down trade-off: coherence and valid novelty rise, soft-preference fit dips slightly, and the breadth tax is a selection-policy effect that separating exploration from commitment recovers. The elastic diagnostic is post-hoc and not part of the primary result.

限制

limitations
  • Controlled synthetic witness only; does not demonstrate that a frontier LLM shows the same effect.

重現

reproduction_instructions
Extract PACC-Hybrid-Lab_v0.1_SYNTHETIC_FINAL.zip; python -m pytest -q; results in results/*.json and docs/PACC_HYBRID_v0.1_RESULTS.md.

結果

來源關係目標狀態ID
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證producesRST-2026-0013 Hybrid v0.1 A/B/C 差值ACTIVEREL-2026-0318

記錄欄位

completed_at
2026-09-09

關係

來源關係目標狀態ID
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證runs_onSYS-2026-0003 PACC-LLM 混合實驗室ACTIVEREL-2026-0313
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證uses_benchmarkBEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkACTIVEREL-2026-0314
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證uses_datasetDAT-2026-0002 PACC-Hybrid v0.1 共享候選池ACTIVEREL-2026-0315
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證testsTHY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限ACTIVEREL-2026-0316
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證producedART-2026-0034 PACC-Hybrid-Lab v0.1 SYNTHETIC FINAL artifact://evemisslab/adaptive-epistemic-systems/PACC-Hybrid-Lab_v0.1_SYNTHETIC_FINAL.zipACTIVEREL-2026-0317
EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證producesRST-2026-0013 Hybrid v0.1 A/B/C 差值ACTIVEREL-2026-0318
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)extendsEXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證ACTIVEREL-2026-0322

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/experiments/EXP-2026-0021/
機器可讀
/ai/experiments/EXP-2026-0021/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at
2026-09-11
generator
tools/extract_aes/extract.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports