理論THY-2026-0106v0.1
結構化品質、hard gate 與規格—驗證分離
品質是 Q(Y | X, S, W, B_Q):輸出相對於任務、規格、環境與邊界。以結構化向量(正確性、對齊、完整、一致、穩健、可驗證、來源)量測,分三層——形式化客觀、結構化半客觀、人類殘餘——致命條件是 hard gate,軟品質不能補償。覆蓋率拆成規格、測試與狀態覆蓋;mutation score 量測試強度;驗證與規格是兩個軸(完美證明了錯的定理什麼都沒解決);評審一致不等於真;成本不是品質,除非規格把它變成品質。品質證據等級從 E(表面有效)到 A+(形式驗證加目標對齊)。
定義
definitions- Q_S = (Q_C, Q_A, Q_K, Q_R, Q_B, Q_V, Q_P), typed per domain; layers Q_L = (Q_F, Q_S, Q_H).
- Hard gate G_H(Y) = ∧ h_i(Y); coverage C_Q = (C_spec, C_test, C_state); mutation score MS = killed / non-equivalent mutants.
- Q_verified = Q_verification ⊗ Q_specification; math vector Q_math = (well-formedness, derivation validity, goal alignment, scope fidelity, axiom transparency, counterexample resistance).
- Quality object 𝔔 = (Q_S, G_H, C_Q, E_Q, Grade_Q, Conf_Q, U_Q, Boundary_Q); scalar Q* = Π_Q(𝔔 | task, projection rule).
前提
assumptions- Evaluation oracles (tests, proof checkers, judge models, humans) are themselves fallible: ObservedQuality = F(TrueQuality, EvaluatorPower, Coverage).
主張
claims- Quality ≠ IntrinsicScalar; SyntacticValidity ≠ SemanticCorrectness; CompileSuccess ≠ CorrectProgram; AllTestsPassed ≠ UniversalCorrectness.
- FormalVerification ≠ RealWorldGoalCorrectness; ProofValidity ≠ GoalEquivalence; ProofGrade ≠ GoalAlignmentGrade.
- FatalConstraintFailure ≁ SoftQualityTradeoff; PeakEpisode ≠ ReliableQuality; Length ≠ Completeness; Cost ≠ Quality.
- ObjectifiableFirst, HumanResidualLast.
形式化
formalization- Robustness sensitivity S_R = ΔQ / d(x, x′); requirement coverage C_R = |satisfied| / n with typed importance.
預測
predictions- Instruments that fold output-format compliance into 'quality' will misreport task competence as failure on tasks the model actually solves.
證偽/失敗條件
falsification_conditions- If a single universal quality scalar predicts downstream task success across domains as well as the structured object does, the structure is redundant.
已知限制
known_limitations- The first pilot showed the prediction in practice — two of three tasks had their 'quality' decided by JSON/Markdown obedience — but that is one instrument on one model, and the theory itself has not been tested beyond it.
證據
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | tests | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0481 |
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | supports | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0489 |
關係
| 來源 | 關係 | 目標 | 狀態 | ID |
|---|---|---|---|---|
RES-2026-0102 品質線——不叫人打分數也能量成果品質 | develops | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0363 |
THY-2026-0107 二元殘餘品質測量(IBQF/BRQM) | extends | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0372 |
PAP-2026-0106 成果品質到底怎麼量?:從形式化正確性到結構化智能品質 | formalizes | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0412 |
EXP-2026-0103 XA-06 第一次真實模型 pilot——Qwythos-9B-v2 走 A0→A5 階梯(36 試驗,2026-09-03) | tests | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0481 |
RST-2026-0102 診斷:三題中兩題量到的「品質」是格式服從性;驗證器序列化失敗 3/18;A3/A4 的上升是缺值假象 | supports | THY-2026-0106 結構化品質、hard gate 與規格—驗證分離 | ACTIVE | REL-2026-0489 |
歷史與來源歷程
- Canonical URL
- https://evemisslab.com/ai/theory/THY-2026-0106/
- 快照
AI-SNAPSHOT-v0.1-fe85b9694a45- 來源歷程
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports