EVEMISSLAB

TheoryTHY-2026-0106v0.1

Structured quality, hard gates and the specification–verification separation

Quality is Q(Y | X, S, W, B_Q): output relative to task, specification, environment and boundary. It is measured as a structured vector (correctness, alignment, completeness, consistency, robustness, verifiability, provenance) in three layers — formal objective, structured semi-objective, human residual — with fatal conditions as a hard gate that soft quality cannot compensate. Coverage splits into specification, test and state coverage; mutation score measures test strength; verification and specification are separate axes (a perfect proof of the wrong theorem solves nothing); evaluator agreement is not truth; cost is not quality unless the specification makes it one. Quality evidence grades run from E (surface validity) to A+ (formal verification plus goal alignment).

Research status
EXPERIMENTAL being tested experimentally
Evidence level
E1 Internal observation
Data basis
THEORY Theoretical reasoning only; no measurement.
Version
0.1
Updated
2026-09-02
Created
2026-09-02
Domain
Evaluation
Program
PRG-2026-0101 Intelligence Physical Metrology (IPM)
Authors
Neo.K (EveMissLab)
AI collaborators
Aletheia (GPT-5.6 Sol, OpenAI ChatGPT)

Definitions

definitions
  • Q_S = (Q_C, Q_A, Q_K, Q_R, Q_B, Q_V, Q_P), typed per domain; layers Q_L = (Q_F, Q_S, Q_H).
  • Hard gate G_H(Y) = ∧ h_i(Y); coverage C_Q = (C_spec, C_test, C_state); mutation score MS = killed / non-equivalent mutants.
  • Q_verified = Q_verification ⊗ Q_specification; math vector Q_math = (well-formedness, derivation validity, goal alignment, scope fidelity, axiom transparency, counterexample resistance).
  • Quality object 𝔔 = (Q_S, G_H, C_Q, E_Q, Grade_Q, Conf_Q, U_Q, Boundary_Q); scalar Q* = Π_Q(𝔔 | task, projection rule).

Assumptions

assumptions
  • Evaluation oracles (tests, proof checkers, judge models, humans) are themselves fallible: ObservedQuality = F(TrueQuality, EvaluatorPower, Coverage).

Claims

claims
  • Quality ≠ IntrinsicScalar; SyntacticValidity ≠ SemanticCorrectness; CompileSuccess ≠ CorrectProgram; AllTestsPassed ≠ UniversalCorrectness.
  • FormalVerification ≠ RealWorldGoalCorrectness; ProofValidity ≠ GoalEquivalence; ProofGrade ≠ GoalAlignmentGrade.
  • FatalConstraintFailure ≁ SoftQualityTradeoff; PeakEpisode ≠ ReliableQuality; Length ≠ Completeness; Cost ≠ Quality.
  • ObjectifiableFirst, HumanResidualLast.

Formalisation

formalization
  • Robustness sensitivity S_R = ΔQ / d(x, x′); requirement coverage C_R = |satisfied| / n with typed importance.

Predictions

predictions
  • Instruments that fold output-format compliance into 'quality' will misreport task competence as failure on tasks the model actually solves.

Falsification / failure conditions

falsification_conditions
  • If a single universal quality scalar predicts downstream task success across domains as well as the structured object does, the structure is redundant.

Known limitations

known_limitations
  • The first pilot showed the prediction in practice — two of three tasks had their 'quality' decided by JSON/Markdown obedience — but that is one instrument on one model, and the theory itself has not been tested beyond it.

Evidence

SourceRelationTargetStatusID
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03)testsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0481
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactsupportsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0489

Relations

SourceRelationTargetStatusID
RES-2026-0102 Quality line — measuring outcome quality without asking humans for a scoredevelopsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0363
THY-2026-0107 Binary residual quality measurement (IBQF / BRQM)extendsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0372
PAP-2026-0106 Paper 06 — How should output quality be measured? From formal correctness to structured intelligence qualityformalizesTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0412
EXP-2026-0103 XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03)testsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0481
RST-2026-0102 Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactsupportsTHY-2026-0106 Structured quality, hard gates and the specification–verification separationACTIVEREL-2026-0489

History and provenance

Canonical URL
https://evemisslab.com/ai/theory/THY-2026-0106/
Machine-readable
/ai/theory/THY-2026-0106/index.json
Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45
Provenance
source
EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by
Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at
2026-09-11
generator
tools/extract_all.py
claim_boundary
status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports