結果RST-2026-0014v0.1
真實本地模型 A/B/C 表(Qwythos-9B-v2)
v0.2 協定第一次在真實語言模型上執行——本地開放權重 9B(Qwythos-9B-v2、Q4_K_M、Ollama、關閉思考)同時當生成器、選擇器與評審;16 題 × 2 次 × 4 候選,380 次呼叫,每條件 32 筆。PACC runtime(C)在 supersession 對齊(相對 B +0.031、相對 A +0.097)與修復成功率(+0.070/+0.094)上升——正是 canonical intent 編譯步驟存在的那兩個軸;衍生一致性(+0.011)與意圖持續(−0.002)沒有動;有效新穎度略低(相對 B −0.045)。多數逐題配對是平手,因為 9B 評審在接近 1.0 處飽和;每題只重複兩次,評審的自由文字 pattern 標籤從不重複,所以創造廣度量不出來。三個條件在 69 % 的題 × 次配對中選了不同的候選。
觀察到的結果
metricsoverallA_llm_onlyhard_adherence- 0.9875
derived_coherence- 0.9812
intent_persistence- 0.9875
supersession_alignment- 0.8094
repair_success- 0.775
usefulness- 0.9528
semantic_novelty- 0.8084
valid_novelty- 0.8475
literal_check_mean- 0.9792
within_task_pattern_entropy_mean- 1.0
B_hard_verifierhard_adherence- 0.9859
derived_coherence- 0.9797
intent_persistence- 0.9875
supersession_alignment- 0.875
repair_success- 0.7984
usefulness- 0.9503
semantic_novelty- 0.7725
valid_novelty- 0.8606
literal_check_mean- 0.9792
within_task_pattern_entropy_mean- 1.0
C_pacc_runtimehard_adherence- 1.0
derived_coherence- 0.9906
intent_persistence- 0.9853
supersession_alignment- 0.9062
repair_success- 0.8688
usefulness- 0.9516
semantic_novelty- 0.7897
valid_novelty- 0.8153
literal_check_mean- 0.9792
within_task_pattern_entropy_mean- 1.0
deltasC-Bhard_adherence- 0.0141
derived_coherence- 0.0109
intent_persistence- -0.0022
supersession_alignment- 0.0312
repair_success- 0.0703
usefulness- 0.0012
semantic_novelty- 0.0172
valid_novelty- -0.0453
literal_check_mean- 0.0
within_task_pattern_entropy_mean- 0.0
C-Ahard_adherence- 0.0125
derived_coherence- 0.0094
intent_persistence- -0.0022
supersession_alignment- 0.0969
repair_success- 0.0938
usefulness- -0.0012
semantic_novelty- -0.0188
valid_novelty- -0.0322
literal_check_mean- 0.0
within_task_pattern_entropy_mean- 0.0
B-Ahard_adherence- -0.0016
derived_coherence- -0.0016
intent_persistence- 0.0
supersession_alignment- 0.0656
repair_success- 0.0234
usefulness- -0.0025
semantic_novelty- -0.0359
valid_novelty- 0.0131
literal_check_mean- 0.0
within_task_pattern_entropy_mean- 0.0
selection_agreementA=B- 0.5625
A=C- 0.40625
B=C- 0.46875
all_same- 0.3125
詮釋
interpretation- Against the v0.2 predeclared interpretation: the predicted coherence and intent-persistence gains over B are not observed; raw novelty did not decrease (semantic novelty +0.017 vs B); the predicted breadth collapse cannot be tested at this repetition count. What did appear — governance gains on supersession and repair with a valid-novelty cost concentrated in multi_constraint and repair tasks — is mechanism-consistent but small, untested statistically, and runs opposite to the synthetic v0.1 valid-novelty picture (+0.196 there). One model, one run, one same-model judge: a first real data point, not a verdict on the architecture.
限制
limitations- Descriptive deltas from one local run with a same-model judge.
主張
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
RST-2026-0014 真實本地模型 A/B/C 表(Qwythos-9B-v2) | qualifies | THY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限 | ACTIVE | REL-2026-0333 |
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
RST-2026-0014 真實本地模型 A/B/C 表(Qwythos-9B-v2) | qualifies | THY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限 | ACTIVE | REL-2026-0333 |
EXP-2026-0023 PACC-Hybrid v0.2——第一次真實模型執行,本地 9B 開放權重模型 | produces | RST-2026-0014 真實本地模型 A/B/C 表(Qwythos-9B-v2) | ACTIVE | REL-2026-0332 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/results/RST-2026-0014/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at- 2026-09-11
generator- tools/extract_aes/extract.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports