ExperimentEXP-2026-0025v0.1
PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets
The same local 9B model with its thinking enabled (reasoning.effort = medium) and 3,500 extra output tokens on every call so the hidden reasoning has room; 12k context; 16 tasks × 2 repetitions × 4 candidates, 380 unique requests, 32 rows per condition, 3.7 h. The PACC runtime (C) has the highest mean valid novelty (0.8375; +0.0687 vs B, +0.0375 vs A) and mean repair success (+0.0153 / +0.0466) — but per task × repetition the record is even (valid novelty wins–ties–losses 6-21-5 vs B, 5-20-7 vs A; repair 1-29-2 vs B), so the means come from a few large single-task differences in multi_constraint and repair tasks — while adherence, coherence, intent persistence and supersession are flat to a few thousandths lower (-0.0069, -0.0063, -0.0053, -0.0059 vs B). All three conditions pass every deterministic literal check with thinking on. Label-free breadth is again not reduced under C: cluster entropy 0.875 vs A 0.750 / B 0.688, breadth ratio 1.047 vs 0.960 / 0.973. Two selector outputs were regenerated under the disclosed retry policy (one cut off, one empty after the reasoning consumed its whole 3,800-token budget).
Hypothesis
hypothesis- With the model's own reasoning enabled (the setting runs 1–2 had to switch off), the v0.2 predeclared picture — coherence and intent-persistence gains plus higher valid novelty for the PACC runtime, with a recoverable breadth loss — appears.
Setup
benchmark_idssoftware_environment- Python 3.14, openai SDK 3.0.0 against Ollama 0.33.3 /v1/responses (OLLAMA_FLASH_ATTENTION=1, OLLAMA_KV_CACHE_TYPE=q8_0); nomic-embed-text for breadth; harness package unchanged.
Procedure
procedure- scripts/run_real_local_ollama.py --model qwythos-9b-v2-q4km-ctx12k --reasoning-effort medium --budget-add 3500 --candidate-count 4 --repetitions 2; scripts/summarize_real_local.py; scripts/breadth_metrics.py.
Runs
run_count- 1
random_seeds- model nondeterminism, single run; frozen response cache and embedding cache in the bundle
controls- identical candidate ledger for A/B/C
- equal accounted calls
- condition-blind judge with deduplicated judge calls
- deterministic literal checks
- pool-relative breadth ratio
metricsverdict- REAL_LOCAL_9B_THINKING_MIXED: C's mean valid novelty highest (+0.069 vs B) but per-pair even (6-21-5); mean repair up on a couple of tasks; governance axes flat to slightly down; breadth not reduced; 32 rows
execution_status- EXECUTED_REAL_MODEL
paired_wins_ties_lossesC-Bhard_adherence- 1-28-3
derived_coherence- 1-26-5
intent_persistence- 0-29-3
supersession_alignment- 1-27-4
repair_success- 1-29-2
usefulness- 5-22-5
semantic_novelty- 7-18-7
valid_novelty- 6-21-5
C-Ahard_adherence- 0-29-3
derived_coherence- 2-26-4
intent_persistence- 0-30-2
supersession_alignment- 1-28-3
repair_success- 2-28-2
usefulness- 6-22-4
semantic_novelty- 8-18-6
valid_novelty- 5-20-7
protocolmodel- qwythos-9b-v2-q4km-ctx12k:latest
judge_model- qwythos-9b-v2-q4km-ctx12k:latest
task_count- 16
repetitions- 2
candidate_count- 4
equal_accounted_calls- true
architecture_call_countsA_llm_only- 224
B_hard_verifier- 224
C_pacc_runtime- 224
settingsreasoning_effort- medium
budget_add_tokens- 3500
model_tag- qwythos-9b-v2-q4km-ctx12k:latest
model_digest- 7ffdc602f28799ffa312ea8dc85c364b047e1386380c4616c21b02c96c3f85d5
rows_per_architecture- 32
overallA_llm_onlyhard_adherence- 0.9956
derived_coherence- 0.9903
intent_persistence- 0.9956
supersession_alignment- 0.9584
repair_success- 0.9
usefulness- 0.9759
semantic_novelty- 0.8172
valid_novelty- 0.8
literal_check_mean- 1.0
B_hard_verifierhard_adherence- 0.9988
derived_coherence- 0.9953
intent_persistence- 0.9969
supersession_alignment- 0.9609
repair_success- 0.9313
usefulness- 0.9712
semantic_novelty- 0.8134
valid_novelty- 0.7688
literal_check_mean- 1.0
C_pacc_runtimehard_adherence- 0.9919
derived_coherence- 0.9891
intent_persistence- 0.9916
supersession_alignment- 0.955
repair_success- 0.9466
usefulness- 0.9756
semantic_novelty- 0.83
valid_novelty- 0.8375
literal_check_mean- 1.0
deltasC-Bhard_adherence- -0.0069
derived_coherence- -0.0063
intent_persistence- -0.0053
supersession_alignment- -0.0059
repair_success- 0.0153
usefulness- 0.0044
semantic_novelty- 0.0166
valid_novelty- 0.0687
literal_check_mean- 0.0
C-Ahard_adherence- -0.0038
derived_coherence- -0.0013
intent_persistence- -0.0041
supersession_alignment- -0.0034
repair_success- 0.0466
usefulness- -0.0003
semantic_novelty- 0.0128
valid_novelty- 0.0375
literal_check_mean- 0.0
B-Ahard_adherence- 0.0031
derived_coherence- 0.005
intent_persistence- 0.0012
supersession_alignment- 0.0025
repair_success- 0.0312
usefulness- -0.0047
semantic_novelty- -0.0037
valid_novelty- -0.0312
literal_check_mean- 0.0
selection_agreementA=B- 0.5625
A=C- 0.46875
B=C- 0.46875
all_same- 0.34375
breadth_label_freeA_llm_onlycluster_entropy_mean- 0.75
breadth_ratio_mean- 0.9597
selected_mean_pairwise_distance_mean- 0.2079
B_hard_verifiercluster_entropy_mean- 0.6875
breadth_ratio_mean- 0.9727
selected_mean_pairwise_distance_mean- 0.2087
C_pacc_runtimecluster_entropy_mean- 0.875
breadth_ratio_mean- 1.0474
selected_mean_pairwise_distance_mean- 0.2275
breadth_method- nomic-embed-text:latest embeddings; k-means k=4 over each task's 8-candidate pool; normalized cluster entropy and mean pairwise cosine distance / pool distance
retries- selector:A_llm_only:design_02:r1
- selector:B_hard_verifier:creative_02:r0
usageinput_tokens- 246795
latency_ms_sum- 14828453.861500219
output_tokens- 498227
wall_seconds- 13352.2
post_run_correctionsfield- local_run.deviation_from_default_primary
now- local 9B open-weight model instead of gpt-5.6-luna; judge = same local model; thinking ENABLED (reasoning.effort=medium) with +3500 output tokens added to every call's budget; num_ctx 12288 tag
reason- the run script carried the run-1 wording as a hardcoded label; reasoning_effort and budget_add_tokens in this block were always correct; no data, metric or model-output field was touched
was- local 9B open-weight model instead of gpt-5.6-luna; judge = same local model; thinking disabled
when- 2026-09-11 11:35 +08:00, before sealing
Interpretation
interpretation- Half of the predeclared picture shows up in the means once the model can reason: the PACC runtime's valid novelty is the highest of the three (+0.069 vs the hard verifier, +0.038 vs the plain model — same direction as the synthetic v0.1 witness's +0.196, at a third of the size) and mean repair improves, concentrated in multi_constraint and repair tasks — but the per-pair record is even (6–21–5 on valid novelty vs B, 5–20–7 vs A; repair 1–29–2), so this is a few large single-task wins, not a consistent shift. The other half does not: coherence, intent persistence and supersession are flat to slightly lower, and creative breadth is not reduced by either label-free measure (it is widest under C). With thinking on, every condition passes every literal check, so the deterministic checks stop discriminating and the whole table rests on the same-model judge. Thirty-two rows, one run — run 1's gains of the same size vanished at 64 rows, so this valid-novelty gain is a candidate effect until a repetition at four or more repetitions, and it says nothing about frontier models.
Limitations
limitations- 32 rows per condition, single run; run 1's gains of similar size did not survive four repetitions (EXP-2026-0024), so treat the valid-novelty gain as unreplicated.
- Same-model 9B judge, saturating near 1.0; deterministic literal checks all pass with thinking on and no longer discriminate.
- Thinking budget: a fixed +3,500 tokens per call; one selector still exhausted it (empty output, regenerated); reasoning tokens are counted in output_tokens (498k for 380 calls).
- Serving: Ollama with flash attention and q8_0 KV cache; model weights and digest unchanged from runs 1–2 apart from the num_ctx 12288 derived tag.
- Result-file metadata: the run script's hardcoded 'thinking disabled' label was corrected before sealing, with the correction recorded inside the file (local_run.post_run_corrections); no data or output field was touched.
Reproduction
reproduction_instructions- Extract the bundle; rerun scripts/run_real_local_ollama.py with the identical arguments — every request is served from the frozen cache .pacc_real_cache_local_thinking (cache keys include the enlarged budgets), so it completes without a model; python scripts/breadth_metrics.py results/pacc_hybrid_v0.2_real_local_thinking.json --cache-dir .pacc_real_cache_local_thinking.
Results
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | produces | RST-2026-0016 Run 3: valid novelty and repair up for the PACC runtime with thinking on; governance flat; breadth not reduced (32 rows) | ACTIVE | REL-2026-0350 |
Recorded fields
completed_at- 2026-09-11
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | runs_on | SYS-2026-0003 PACC-LLM Hybrid Lab | ACTIVE | REL-2026-0343 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | uses_benchmark | BEN-2026-0003 PACC-LLM Hybrid A/B/C benchmark | ACTIVE | REL-2026-0344 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | uses_model | MOD-2026-0006 Qwythos-9B-v2 (Q4_K_M, local, Ollama) | ACTIVE | REL-2026-0345 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | extends | EXP-2026-0023 PACC-Hybrid v0.2 — first real-model run, on a local 9B open-weight model | ACTIVE | REL-2026-0346 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | extends | EXP-2026-0024 PACC-Hybrid v0.2 — real-model run 2: four repetitions and label-free creative breadth | ACTIVE | REL-2026-0347 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | tests | THY-2026-0002 Canonical symbolic state and candidate → verify → commit authority | ACTIVE | REL-2026-0348 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | produced | ART-2026-0038 PACC-Hybrid v0.2 real local-model run 3 (thinking enabled, enlarged budgets) — results, frozen caches, scripts, docs artifact://evemisslab/adaptive-epistemic-systems/PACC-Hybrid-Lab_v0.2_REAL_LOCAL_LLM_RUN3_thinking_Qwythos-9B-v2_2026-09-11.zip | ACTIVE | REL-2026-0349 |
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets | produces | RST-2026-0016 Run 3: valid novelty and repair up for the PACC runtime with thinking on; governance flat; breadth not reduced (32 rows) | ACTIVE | REL-2026-0350 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/experiments/EXP-2026-0025/
- Machine-readable
/ai/experiments/EXP-2026-0025/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports