實驗EXP-2026-0025v0.1
PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算
同一顆本地 9B 模型,開啟思考(reasoning.effort = medium),每次呼叫多給 3,500 個輸出 token 讓隱藏推理有空間;12k context;16 題 × 2 次 × 4 候選,380 次唯一請求,每條件 32 筆,3.7 小時。PACC runtime(C)的平均有效新穎度最高(0.8375;相對 B +0.0687、相對 A +0.0375),平均修復成功率也最高(+0.0153/+0.0466)——但逐題 × 次配對的戰績是平的(有效新穎度勝–平–負:對 B 6-21-5、對 A 5-20-7;修復對 B 1-29-2),平均差來自 multi_constraint 與 repair 題裡少數幾個大差距——遵守、一致性、意圖持續與 supersession 則持平到低幾個千分點(相對 B -0.0069、-0.0063、-0.0053、-0.0059)。開啟思考後三個條件全部通過所有確定性字面檢查。標籤無關廣度在 C 下再次沒有縮減:群熵 0.875,A 0.750/B 0.688;廣度比 1.047,對 0.960/0.973。兩次 selector 輸出依公開的重試政策重新生成(一次被截斷、一次推理吃光 3,800 token 預算後輸出為空)。
假設
hypothesis- With the model's own reasoning enabled (the setting runs 1–2 had to switch off), the v0.2 predeclared picture — coherence and intent-persistence gains plus higher valid novelty for the PACC runtime, with a recoverable breadth loss — appears.
設定
benchmark_idssoftware_environment- Python 3.14, openai SDK 3.0.0 against Ollama 0.33.3 /v1/responses (OLLAMA_FLASH_ATTENTION=1, OLLAMA_KV_CACHE_TYPE=q8_0); nomic-embed-text for breadth; harness package unchanged.
程序
procedure- scripts/run_real_local_ollama.py --model qwythos-9b-v2-q4km-ctx12k --reasoning-effort medium --budget-add 3500 --candidate-count 4 --repetitions 2; scripts/summarize_real_local.py; scripts/breadth_metrics.py.
執行
run_count- 1
random_seeds- model nondeterminism, single run; frozen response cache and embedding cache in the bundle
controls- identical candidate ledger for A/B/C
- equal accounted calls
- condition-blind judge with deduplicated judge calls
- deterministic literal checks
- pool-relative breadth ratio
metricsverdict- REAL_LOCAL_9B_THINKING_MIXED: C's mean valid novelty highest (+0.069 vs B) but per-pair even (6-21-5); mean repair up on a couple of tasks; governance axes flat to slightly down; breadth not reduced; 32 rows
execution_status- EXECUTED_REAL_MODEL
paired_wins_ties_lossesC-Bhard_adherence- 1-28-3
derived_coherence- 1-26-5
intent_persistence- 0-29-3
supersession_alignment- 1-27-4
repair_success- 1-29-2
usefulness- 5-22-5
semantic_novelty- 7-18-7
valid_novelty- 6-21-5
C-Ahard_adherence- 0-29-3
derived_coherence- 2-26-4
intent_persistence- 0-30-2
supersession_alignment- 1-28-3
repair_success- 2-28-2
usefulness- 6-22-4
semantic_novelty- 8-18-6
valid_novelty- 5-20-7
protocolmodel- qwythos-9b-v2-q4km-ctx12k:latest
judge_model- qwythos-9b-v2-q4km-ctx12k:latest
task_count- 16
repetitions- 2
candidate_count- 4
equal_accounted_calls- true
architecture_call_countsA_llm_only- 224
B_hard_verifier- 224
C_pacc_runtime- 224
settingsreasoning_effort- medium
budget_add_tokens- 3500
model_tag- qwythos-9b-v2-q4km-ctx12k:latest
model_digest- 7ffdc602f28799ffa312ea8dc85c364b047e1386380c4616c21b02c96c3f85d5
rows_per_architecture- 32
overallA_llm_onlyhard_adherence- 0.9956
derived_coherence- 0.9903
intent_persistence- 0.9956
supersession_alignment- 0.9584
repair_success- 0.9
usefulness- 0.9759
semantic_novelty- 0.8172
valid_novelty- 0.8
literal_check_mean- 1.0
B_hard_verifierhard_adherence- 0.9988
derived_coherence- 0.9953
intent_persistence- 0.9969
supersession_alignment- 0.9609
repair_success- 0.9313
usefulness- 0.9712
semantic_novelty- 0.8134
valid_novelty- 0.7688
literal_check_mean- 1.0
C_pacc_runtimehard_adherence- 0.9919
derived_coherence- 0.9891
intent_persistence- 0.9916
supersession_alignment- 0.955
repair_success- 0.9466
usefulness- 0.9756
semantic_novelty- 0.83
valid_novelty- 0.8375
literal_check_mean- 1.0
deltasC-Bhard_adherence- -0.0069
derived_coherence- -0.0063
intent_persistence- -0.0053
supersession_alignment- -0.0059
repair_success- 0.0153
usefulness- 0.0044
semantic_novelty- 0.0166
valid_novelty- 0.0687
literal_check_mean- 0.0
C-Ahard_adherence- -0.0038
derived_coherence- -0.0013
intent_persistence- -0.0041
supersession_alignment- -0.0034
repair_success- 0.0466
usefulness- -0.0003
semantic_novelty- 0.0128
valid_novelty- 0.0375
literal_check_mean- 0.0
B-Ahard_adherence- 0.0031
derived_coherence- 0.005
intent_persistence- 0.0012
supersession_alignment- 0.0025
repair_success- 0.0312
usefulness- -0.0047
semantic_novelty- -0.0037
valid_novelty- -0.0312
literal_check_mean- 0.0
selection_agreementA=B- 0.5625
A=C- 0.46875
B=C- 0.46875
all_same- 0.34375
breadth_label_freeA_llm_onlycluster_entropy_mean- 0.75
breadth_ratio_mean- 0.9597
selected_mean_pairwise_distance_mean- 0.2079
B_hard_verifiercluster_entropy_mean- 0.6875
breadth_ratio_mean- 0.9727
selected_mean_pairwise_distance_mean- 0.2087
C_pacc_runtimecluster_entropy_mean- 0.875
breadth_ratio_mean- 1.0474
selected_mean_pairwise_distance_mean- 0.2275
breadth_method- nomic-embed-text:latest embeddings; k-means k=4 over each task's 8-candidate pool; normalized cluster entropy and mean pairwise cosine distance / pool distance
retries- selector:A_llm_only:design_02:r1
- selector:B_hard_verifier:creative_02:r0
usageinput_tokens- 246795
latency_ms_sum- 14828453.861500219
output_tokens- 498227
wall_seconds- 13352.2
post_run_correctionsfield- local_run.deviation_from_default_primary
now- local 9B open-weight model instead of gpt-5.6-luna; judge = same local model; thinking ENABLED (reasoning.effort=medium) with +3500 output tokens added to every call's budget; num_ctx 12288 tag
reason- the run script carried the run-1 wording as a hardcoded label; reasoning_effort and budget_add_tokens in this block were always correct; no data, metric or model-output field was touched
was- local 9B open-weight model instead of gpt-5.6-luna; judge = same local model; thinking disabled
when- 2026-09-11 11:35 +08:00, before sealing
詮釋
interpretation- Half of the predeclared picture shows up in the means once the model can reason: the PACC runtime's valid novelty is the highest of the three (+0.069 vs the hard verifier, +0.038 vs the plain model — same direction as the synthetic v0.1 witness's +0.196, at a third of the size) and mean repair improves, concentrated in multi_constraint and repair tasks — but the per-pair record is even (6–21–5 on valid novelty vs B, 5–20–7 vs A; repair 1–29–2), so this is a few large single-task wins, not a consistent shift. The other half does not: coherence, intent persistence and supersession are flat to slightly lower, and creative breadth is not reduced by either label-free measure (it is widest under C). With thinking on, every condition passes every literal check, so the deterministic checks stop discriminating and the whole table rests on the same-model judge. Thirty-two rows, one run — run 1's gains of the same size vanished at 64 rows, so this valid-novelty gain is a candidate effect until a repetition at four or more repetitions, and it says nothing about frontier models.
限制
limitations- 32 rows per condition, single run; run 1's gains of similar size did not survive four repetitions (EXP-2026-0024), so treat the valid-novelty gain as unreplicated.
- Same-model 9B judge, saturating near 1.0; deterministic literal checks all pass with thinking on and no longer discriminate.
- Thinking budget: a fixed +3,500 tokens per call; one selector still exhausted it (empty output, regenerated); reasoning tokens are counted in output_tokens (498k for 380 calls).
- Serving: Ollama with flash attention and q8_0 KV cache; model weights and digest unchanged from runs 1–2 apart from the num_ctx 12288 derived tag.
- Result-file metadata: the run script's hardcoded 'thinking disabled' label was corrected before sealing, with the correction recorded inside the file (local_run.post_run_corrections); no data or output field was touched.
重現
reproduction_instructions- Extract the bundle; rerun scripts/run_real_local_ollama.py with the identical arguments — every request is served from the frozen cache .pacc_real_cache_local_thinking (cache keys include the enlarged budgets), so it completes without a model; python scripts/breadth_metrics.py results/pacc_hybrid_v0.2_real_local_thinking.json --cache-dir .pacc_real_cache_local_thinking.
結果
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | produces | RST-2026-0016 第三次執行:開啟思考後 PACC runtime 的有效新穎度與修復上升;治理持平;廣度未縮減(32 筆) | ACTIVE | REL-2026-0350 |
記錄欄位
completed_at- 2026-09-11
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | runs_on | SYS-2026-0003 PACC-LLM 混合實驗室 | ACTIVE | REL-2026-0343 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | uses_benchmark | BEN-2026-0003 PACC-LLM 混合 A/B/C benchmark | ACTIVE | REL-2026-0344 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | uses_model | MOD-2026-0006 Qwythos-9B-v2(Q4_K_M,本地,Ollama) | ACTIVE | REL-2026-0345 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | extends | EXP-2026-0023 PACC-Hybrid v0.2——第一次真實模型執行,本地 9B 開放權重模型 | ACTIVE | REL-2026-0346 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | extends | EXP-2026-0024 PACC-Hybrid v0.2——真實模型第二次執行:四次重複與標籤無關的創造廣度 | ACTIVE | REL-2026-0347 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | tests | THY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限 | ACTIVE | REL-2026-0348 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | produced | ART-2026-0038 PACC-Hybrid v0.2 real local-model run 3 (thinking enabled, enlarged budgets) — results, frozen caches, scripts, docs artifact://evemisslab/adaptive-epistemic-systems/PACC-Hybrid-Lab_v0.2_REAL_LOCAL_LLM_RUN3_thinking_Qwythos-9B-v2_2026-09-11.zip | ACTIVE | REL-2026-0349 |
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算 | produces | RST-2026-0016 第三次執行:開啟思考後 PACC runtime 的有效新穎度與修復上升;治理持平;廣度未縮減(32 筆) | ACTIVE | REL-2026-0350 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0025/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports