EVEMISSLAB

ExperimentEXP-2026-0025v0.1

PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgets

The same local 9B model with its thinking enabled (reasoning.effort = medium) and 3,500 extra output tokens on every call so the hidden reasoning has room; 12k context; 16 tasks × 2 repetitions × 4 candidates, 380 unique requests, 32 rows per condition, 3.7 h. The PACC runtime (C) has the highest mean valid novelty (0.8375; +0.0687 vs B, +0.0375 vs A) and mean repair success (+0.0153 / +0.0466) — but per task × repetition the record is even (valid novelty wins–ties–losses 6-21-5 vs B, 5-20-7 vs A; repair 1-29-2 vs B), so the means come from a few large single-task differences in multi_constraint and repair tasks — while adherence, coherence, intent persistence and supersession are flat to a few thousandths lower (-0.0069, -0.0063, -0.0053, -0.0059 vs B). All three conditions pass every deterministic literal check with thinking on. Label-free breadth is again not reduced under C: cluster entropy 0.875 vs A 0.750 / B 0.688, breadth ratio 1.047 vs 0.960 / 0.973. Two selector outputs were regenerated under the disclosed retry policy (one cut off, one empty after the reasoning consumed its whole 3,800-token budget).

Research status
STABLE the current conclusions are relatively stable
Evidence level
E2 Controlled experiment
Result
MIXED
Data basis
REAL MODEL A real language model was executed; the model, its version and its configuration are recorded on the page.
Version
0.1
Updated
2026-09-11
Created
2026-09-11
Domain
Reasoning, Evaluation
Program
PRG-2026-0001 Adaptive Epistemic Systems
Authors
Neo.K (EveMissLab)
AI collaborators
Sol (GPT-5.6, OpenAI ChatGPT)

Hypothesis

hypothesis
With the model's own reasoning enabled (the setting runs 1–2 had to switch off), the v0.2 predeclared picture — coherence and intent-persistence gains plus higher valid novelty for the PACC runtime, with a recoverable breadth loss — appears.

Setup

model_ids
benchmark_ids
software_environment
Python 3.14, openai SDK 3.0.0 against Ollama 0.33.3 /v1/responses (OLLAMA_FLASH_ATTENTION=1, OLLAMA_KV_CACHE_TYPE=q8_0); nomic-embed-text for breadth; harness package unchanged.

Procedure

procedure
scripts/run_real_local_ollama.py --model qwythos-9b-v2-q4km-ctx12k --reasoning-effort medium --budget-add 3500 --candidate-count 4 --repetitions 2; scripts/summarize_real_local.py; scripts/breadth_metrics.py.

Runs

run_count
1
random_seeds
  • model nondeterminism, single run; frozen response cache and embedding cache in the bundle
controls
  • identical candidate ledger for A/B/C
  • equal accounted calls
  • condition-blind judge with deduplicated judge calls
  • deterministic literal checks
  • pool-relative breadth ratio
metrics
verdict
REAL_LOCAL_9B_THINKING_MIXED: C's mean valid novelty highest (+0.069 vs B) but per-pair even (6-21-5); mean repair up on a couple of tasks; governance axes flat to slightly down; breadth not reduced; 32 rows
execution_status
EXECUTED_REAL_MODEL
paired_wins_ties_losses
C-B
hard_adherence
1-28-3
derived_coherence
1-26-5
intent_persistence
0-29-3
supersession_alignment
1-27-4
repair_success
1-29-2
usefulness
5-22-5
semantic_novelty
7-18-7
valid_novelty
6-21-5
C-A
hard_adherence
0-29-3
derived_coherence
2-26-4
intent_persistence
0-30-2
supersession_alignment
1-28-3
repair_success
2-28-2
usefulness
6-22-4
semantic_novelty
8-18-6
valid_novelty
5-20-7
protocol
model
qwythos-9b-v2-q4km-ctx12k:latest
judge_model
qwythos-9b-v2-q4km-ctx12k:latest
task_count
16
repetitions
2
candidate_count
4
equal_accounted_calls
true
architecture_call_counts
A_llm_only
224
B_hard_verifier
224
C_pacc_runtime
224
settings
reasoning_effort
medium
budget_add_tokens
3500
model_tag
qwythos-9b-v2-q4km-ctx12k:latest
model_digest
7ffdc602f28799ffa312ea8dc85c364b047e1386380c4616c21b02c96c3f85d5
rows_per_architecture
32
overall
A_llm_only
hard_adherence
0.9956
derived_coherence
0.9903
intent_persistence
0.9956
supersession_alignment
0.9584
repair_success
0.9
usefulness
0.9759
semantic_novelty
0.8172
valid_novelty
0.8
literal_check_mean
1.0
B_hard_verifier
hard_adherence
0.9988
derived_coherence
0.9953
intent_persistence
0.9969
supersession_alignment
0.9609
repair_success
0.9313
usefulness
0.9712
semantic_novelty
0.8134
valid_novelty
0.7688
literal_check_mean
1.0
C_pacc_runtime
hard_adherence
0.9919
derived_coherence
0.9891
intent_persistence
0.9916
supersession_alignment
0.955
repair_success
0.9466
usefulness
0.9756
semantic_novelty
0.83
valid_novelty
0.8375
literal_check_mean
1.0
deltas
C-B
hard_adherence
-0.0069
derived_coherence
-0.0063
intent_persistence
-0.0053
supersession_alignment
-0.0059
repair_success
0.0153
usefulness
0.0044
semantic_novelty
0.0166
valid_novelty
0.0687
literal_check_mean
0.0
C-A
hard_adherence
-0.0038
derived_coherence
-0.0013
intent_persistence
-0.0041
supersession_alignment
-0.0034
repair_success
0.0466
usefulness
-0.0003
semantic_novelty
0.0128
valid_novelty
0.0375
literal_check_mean
0.0
B-A
hard_adherence
0.0031
derived_coherence
0.005
intent_persistence
0.0012
supersession_alignment
0.0025
repair_success
0.0312
usefulness
-0.0047
semantic_novelty
-0.0037
valid_novelty
-0.0312
literal_check_mean
0.0
selection_agreement
A=B
0.5625
A=C
0.46875
B=C
0.46875
all_same
0.34375
breadth_label_free
A_llm_only
cluster_entropy_mean
0.75
breadth_ratio_mean
0.9597
selected_mean_pairwise_distance_mean
0.2079
B_hard_verifier
cluster_entropy_mean
0.6875
breadth_ratio_mean
0.9727
selected_mean_pairwise_distance_mean
0.2087
C_pacc_runtime
cluster_entropy_mean
0.875
breadth_ratio_mean
1.0474
selected_mean_pairwise_distance_mean
0.2275
breadth_method
nomic-embed-text:latest embeddings; k-means k=4 over each task's 8-candidate pool; normalized cluster entropy and mean pairwise cosine distance / pool distance
retries
  • selector:A_llm_only:design_02:r1
  • selector:B_hard_verifier:creative_02:r0
usage
input_tokens
246795
latency_ms_sum
14828453.861500219
output_tokens
498227
wall_seconds
13352.2
post_run_corrections
  • field
    local_run.deviation_from_default_primary
    now
    local 9B open-weight model instead of gpt-5.6-luna; judge = same local model; thinking ENABLED (reasoning.effort=medium) with +3500 output tokens added to every call's budget; num_ctx 12288 tag
    reason
    the run script carried the run-1 wording as a hardcoded label; reasoning_effort and budget_add_tokens in this block were always correct; no data, metric or model-output field was touched
    was
    local 9B open-weight model instead of gpt-5.6-luna; judge = same local model; thinking disabled
    when
    2026-09-11 11:35 +08:00, before sealing

Interpretation

interpretation
Half of the predeclared picture shows up in the means once the model can reason: the PACC runtime's valid novelty is the highest of the three (+0.069 vs the hard verifier, +0.038 vs the plain model — same direction as the synthetic v0.1 witness's +0.196, at a third of the size) and mean repair improves, concentrated in multi_constraint and repair tasks — but the per-pair record is even (6–21–5 on valid novelty vs B, 5–20–7 vs A; repair 1–29–2), so this is a few large single-task wins, not a consistent shift. The other half does not: coherence, intent persistence and supersession are flat to slightly lower, and creative breadth is not reduced by either label-free measure (it is widest under C). With thinking on, every condition passes every literal check, so the deterministic checks stop discriminating and the whole table rests on the same-model judge. Thirty-two rows, one run — run 1's gains of the same size vanished at 64 rows, so this valid-novelty gain is a candidate effect until a repetition at four or more repetitions, and it says nothing about frontier models.

Limitations

limitations
  • 32 rows per condition, single run; run 1's gains of similar size did not survive four repetitions (EXP-2026-0024), so treat the valid-novelty gain as unreplicated.
  • Same-model 9B judge, saturating near 1.0; deterministic literal checks all pass with thinking on and no longer discriminate.
  • Thinking budget: a fixed +3,500 tokens per call; one selector still exhausted it (empty output, regenerated); reasoning tokens are counted in output_tokens (498k for 380 calls).
  • Serving: Ollama with flash attention and q8_0 KV cache; model weights and digest unchanged from runs 1–2 apart from the num_ctx 12288 derived tag.
  • Result-file metadata: the run script's hardcoded 'thinking disabled' label was corrected before sealing, with the correction recorded inside the file (local_run.post_run_corrections); no data or output field was touched.

Reproduction

reproduction_instructions
Extract the bundle; rerun scripts/run_real_local_ollama.py with the identical arguments — every request is served from the frozen cache .pacc_real_cache_local_thinking (cache keys include the enlarged budgets), so it completes without a model; python scripts/breadth_metrics.py results/pacc_hybrid_v0.2_real_local_thinking.json --cache-dir .pacc_real_cache_local_thinking.

Results

SourceRelationTargetStatusID
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsproducesRST-2026-0016 Run 3: valid novelty and repair up for the PACC runtime with thinking on; governance flat; breadth not reduced (32 rows)ACTIVEREL-2026-0350

Recorded fields

completed_at
2026-09-11

Relations

SourceRelationTargetStatusID
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsruns_onSYS-2026-0003 PACC-LLM Hybrid LabACTIVEREL-2026-0343
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsuses_benchmarkBEN-2026-0003 PACC-LLM Hybrid A/B/C benchmarkACTIVEREL-2026-0344
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsuses_modelMOD-2026-0006 Qwythos-9B-v2 (Q4_K_M, local, Ollama)ACTIVEREL-2026-0345
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsextendsEXP-2026-0023 PACC-Hybrid v0.2 — first real-model run, on a local 9B open-weight modelACTIVEREL-2026-0346
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsextendsEXP-2026-0024 PACC-Hybrid v0.2 — real-model run 2: four repetitions and label-free creative breadthACTIVEREL-2026-0347
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetstestsTHY-2026-0002 Canonical symbolic state and candidate → verify → commit authorityACTIVEREL-2026-0348
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsproducedART-2026-0038 PACC-Hybrid v0.2 real local-model run 3 (thinking enabled, enlarged budgets) — results, frozen caches, scripts, docs artifact://evemisslab/adaptive-epistemic-systems/PACC-Hybrid-Lab_v0.2_REAL_LOCAL_LLM_RUN3_thinking_Qwythos-9B-v2_2026-09-11.zipACTIVEREL-2026-0349
EXP-2026-0025 PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsproducesRST-2026-0016 Run 3: valid novelty and repair up for the PACC runtime with thinking on; governance flat; breadth not reduced (32 rows)ACTIVEREL-2026-0350

History and provenance

Canonical URL
https://evemisslab.com/ai/experiments/EXP-2026-0025/
Machine-readable
/ai/experiments/EXP-2026-0025/index.json
Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45
Provenance
source
EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at
2026-09-11
generator
tools/extract_aes/extract.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports