EVEMISSLAB
English

結果RST-2026-0016v0.1

第三次執行:開啟思考後 PACC runtime 的有效新穎度與修復上升;治理持平;廣度未縮減(32 筆)

C 相對 B:有效新穎度 +0.0687、修復 +0.0153、語義新穎度 +0.0166、衍生一致性 -0.0063、意圖 -0.0053、supersession -0.0059;C 相對 A:有效新穎度 +0.0375、修復 +0.0466;標籤無關廣度比 C 1.047/A 0.960/B 0.973。

研究狀態
STABLE 目前的研究結論相對穩定
證據等級
E3 重複實驗
結果
MIXED
資料基礎
REAL MODEL 真的跑了語言模型;模型、版本與設定都記在頁面上。
版本
0.1
更新
2026-09-11
建立
2026-09-11
領域
Evaluation
計畫
PRG-2026-0001 自適應世界狀態系統的第一原理框架
作者
Neo.K (EveMissLab)
AI 協作
Sol (GPT-5.6, OpenAI ChatGPT)

觀察到的結果

metrics
deltas
C-B
hard_adherence
-0.0069
derived_coherence
-0.0063
intent_persistence
-0.0053
supersession_alignment
-0.0059
repair_success
0.0153
usefulness
0.0044
semantic_novelty
0.0166
valid_novelty
0.0687
literal_check_mean
0.0
C-A
hard_adherence
-0.0038
derived_coherence
-0.0013
intent_persistence
-0.0041
supersession_alignment
-0.0034
repair_success
0.0466
usefulness
-0.0003
semantic_novelty
0.0128
valid_novelty
0.0375
literal_check_mean
0.0
B-A
hard_adherence
0.0031
derived_coherence
0.005
intent_persistence
0.0012
supersession_alignment
0.0025
repair_success
0.0312
usefulness
-0.0047
semantic_novelty
-0.0037
valid_novelty
-0.0312
literal_check_mean
0.0
paired_wins_ties_losses
C-B
hard_adherence
1-28-3
derived_coherence
1-26-5
intent_persistence
0-29-3
supersession_alignment
1-27-4
repair_success
1-29-2
usefulness
5-22-5
semantic_novelty
7-18-7
valid_novelty
6-21-5
C-A
hard_adherence
0-29-3
derived_coherence
2-26-4
intent_persistence
0-30-2
supersession_alignment
1-28-3
repair_success
2-28-2
usefulness
6-22-4
semantic_novelty
8-18-6
valid_novelty
5-20-7
breadth_label_free
A_llm_only
cluster_entropy_mean
0.75
breadth_ratio_mean
0.9597
selected_mean_pairwise_distance_mean
0.2079
B_hard_verifier
cluster_entropy_mean
0.6875
breadth_ratio_mean
0.9727
selected_mean_pairwise_distance_mean
0.2087
C_pacc_runtime
cluster_entropy_mean
0.875
breadth_ratio_mean
1.0474
selected_mean_pairwise_distance_mean
0.2275
selection_agreement
A=B
0.5625
A=C
0.46875
B=C
0.46875
all_same
0.34375

詮釋

interpretation
The valid-novelty half of the synthetic prediction appears in the means, at a third of the synthetic size, once the model reasons — carried by a few large single-task wins (per-pair 6–21–5 vs B); the coherence/intent half and the breadth-loss prediction do not appear. Unreplicated: a same-size gain in run 1 vanished at 64 rows.

限制

limitations
  • Descriptive; 32 rows; same-model judge; one model family; single thinking budget.

主張

來源關係目標狀態ID
RST-2026-0016 第三次執行:開啟思考後 PACC runtime 的有效新穎度與修復上升;治理持平;廣度未縮減(32 筆)qualifiesTHY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限ACTIVEREL-2026-0351

關係

來源關係目標狀態ID
RST-2026-0016 第三次執行:開啟思考後 PACC runtime 的有效新穎度與修復上升;治理持平;廣度未縮減(32 筆)qualifiesTHY-2026-0002 Canonical 符號狀態與 candidate → verify → commit 權限ACTIVEREL-2026-0351
EXP-2026-0025 PACC-Hybrid v0.2——真實模型第三次執行:開啟思考並放大輸出預算producesRST-2026-0016 第三次執行:開啟思考後 PACC runtime 的有效新穎度與修復上升;治理持平;廣度未縮減(32 筆)ACTIVEREL-2026-0350

歷史與來源歷程

Canonical URL
https://evemisslab.com/ai/results/RST-2026-0016/
機器可讀
/ai/results/RST-2026-0016/index.json
快照
AI-SNAPSHOT-v0.1-fe85b9694a45
來源歷程
source
EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
extracted_at
2026-09-11
generator
tools/extract_aes/extract.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports