實驗EXP-2026-0022v0.1
PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)
16 個橫跨預登記家族的手寫自然語言任務;A/B/C 共用一本候選帳本與相同計費預算;C 拿不到 gold 約束或 supersession metadata;gold rubric 只有對條件盲的評審看得到;字面機器檢查獨立於評審;相同答案重用同一個評審快取鍵;沒有 OPENAI_API_KEY 時 fail-closed;凍結快取重播提供者讓 live run 後可精確重播。25 個測試通過。執行環境沒有 API key,因此不存在任何真實模型輸出;隨附的 mock smoke 檔標為 MOCK_ONLY_NOT_REAL_MODEL。
假設
hypothesis- A real language model under the PACC runtime shows the coherence / valid-novelty gains and recoverable breadth loss seen in the synthetic witness.
設定
benchmark_idssoftware_environment- Python; OpenAI API provider (fails closed without key); deterministic fake provider for protocol tests only.
執行
run_count- 0
metricsverdict- REAL_LLM_HARNESS_VALIDATED_BUT_REAL_MODEL_NOT_EXECUTED
execution_status- NOT_EXECUTED_REAL_MODEL
tests_passed- 25
tasks- 16
real_llm_calls- 0
詮釋
interpretation- Closes the harness, not the scientific question. Next action: a small real smoke (6 tasks × 2 repetitions × 3 candidates), freeze the cache, inspect blind-evaluator consistency, then the full 16-task primary without changing prompts or metrics.
限制
limitations- No claim about real-model reasoning, intent understanding, imagination, human-rated usefulness, cross-model transfer or hallucination is permitted before a real-model result exists.
- A single-model judge is not human evaluation even after a live run.
- Executed for the first time on 2026-09-11 with a local open-weight model — see EXP-2026-0023.
重現
reproduction_instructions- Extract PACC-Hybrid-Lab_v0.2_REAL_LLM_HARNESS_FINAL.zip; python -m pytest -q (25 tests); set OPENAI_API_KEY and run the smoke per docs/REPRODUCIBILITY_v0.2.md.
記錄欄位
random_seeds
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行) | runs_on | SYS-2026-0003 PACC-LLM 混合實驗室 | ACTIVE | REL-2026-0320 |
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行) | uses_benchmark | BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | ACTIVE | REL-2026-0321 |
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行) | extends | EXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證 | ACTIVE | REL-2026-0322 |
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行) | tests | THY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限 | ACTIVE | REL-2026-0323 |
EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行) | produced | ART-2026-0035 PACC-Hybrid-Lab v0.2 REAL LLM HARNESS FINAL artifact://evemisslab/adaptive-epistemic-systems/PACC-Hybrid-Lab_v0.2_REAL_LLM_HARNESS_FINAL.zip | ACTIVE | REL-2026-0324 |
EXP-2026-0023 PACC-Hybrid v0.2——第一次真實模型執行,本地 9B 開放權重模型 | extends | EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行) | ACTIVE | REL-2026-0329 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0022/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports