EVEMISSLAB
English

實驗EXP-2026-0022v0.1

PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)

16 個橫跨預登記家族的手寫自然語言任務;A/B/C 共用一本候選帳本與相同計費預算;C 拿不到 gold 約束或 supersession metadata;gold rubric 只有對條件盲的評審看得到;字面機器檢查獨立於評審;相同答案重用同一個評審快取鍵;沒有 OPENAI_API_KEY 時 fail-closed;凍結快取重播提供者讓 live run 後可精確重播。25 個測試通過。執行環境沒有 API key,因此不存在任何真實模型輸出;隨附的 mock smoke 檔標為 MOCK_ONLY_NOT_REAL_MODEL。

研究狀態
STABLE 目前的研究結論相對穩定
證據等級
E1 內部觀察
結果
INCONCLUSIVE
資料基礎
NOT RUN 只設計並驗證了 harness,從未用真實模型執行。
版本
0.1
更新
2026-09-09
建立
2026-09-09
領域
Reasoning, Evaluation
計畫
PRG-2026-0001 自適應世界狀態系統的第一原理框架
作者
Neo.K (EveMissLab)
AI 協作
Sol (GPT-5.6, OpenAI ChatGPT)

假設

hypothesis
A real language model under the PACC runtime shows the coherence / valid-novelty gains and recoverable breadth loss seen in the synthetic witness.

設定

benchmark_ids
software_environment
Python; OpenAI API provider (fails closed without key); deterministic fake provider for protocol tests only.

執行

run_count
0
metrics
verdict
REAL_LLM_HARNESS_VALIDATED_BUT_REAL_MODEL_NOT_EXECUTED
execution_status
NOT_EXECUTED_REAL_MODEL
tests_passed
25
tasks
16
real_llm_calls
0

詮釋

interpretation
Closes the harness, not the scientific question. Next action: a small real smoke (6 tasks × 2 repetitions × 3 candidates), freeze the cache, inspect blind-evaluator consistency, then the full 16-task primary without changing prompts or metrics.

限制

limitations
  • No claim about real-model reasoning, intent understanding, imagination, human-rated usefulness, cross-model transfer or hallucination is permitted before a real-model result exists.
  • A single-model judge is not human evaluation even after a live run.
  • Executed for the first time on 2026-09-11 with a local open-weight model — see EXP-2026-0023.

重現

reproduction_instructions
Extract PACC-Hybrid-Lab_v0.2_REAL_LLM_HARNESS_FINAL.zip; python -m pytest -q (25 tests); set OPENAI_API_KEY and run the smoke per docs/REPRODUCIBILITY_v0.2.md.

記錄欄位

random_seeds

    關係

    來源關係目標狀態ID
    EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)runs_onSYS-2026-0003 PACC-LLM 混合實驗室ACTIVEREL-2026-0320
    EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)uses_benchmarkBEN-2026-0003 PACC-LLM 混合 A/B/C benchmarkACTIVEREL-2026-0321
    EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)extendsEXP-2026-0021 PACC-Hybrid v0.1——合成 A/B/C 架構見證ACTIVEREL-2026-0322
    EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)testsTHY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限ACTIVEREL-2026-0323
    EXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)producedART-2026-0035 PACC-Hybrid-Lab v0.2 REAL LLM HARNESS FINAL artifact://evemisslab/adaptive-epistemic-systems/PACC-Hybrid-Lab_v0.2_REAL_LLM_HARNESS_FINAL.zipACTIVEREL-2026-0324
    EXP-2026-0023 PACC-Hybrid v0.2——第一次真實模型執行,本地 9B 開放權重模型extendsEXP-2026-0022 PACC-Hybrid v0.2——真實語言模型 A/B/C harness(尚未執行)ACTIVEREL-2026-0329

    歷史與來源歷程

    Canonical URL
    https://evemisslab.com/ai/experiments/EXP-2026-0022/
    機器可讀
    /ai/experiments/EXP-2026-0022/index.json
    快照
    AI-SNAPSHOT-v0.1-fe85b9694a45
    來源歷程
    source
    EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
    extracted_by
    Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
    extracted_at
    2026-09-11
    generator
    tools/extract_aes/extract.py
    claim_boundary
    status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports