TheoryTHY-2026-0108v0.1
Typed, versioned quality ontology for high-ambiguity artifacts
Before any item is written, quality must be a typed space Q[domain, task, context, audience] = Q_core ⊕ Q_domain ⊕ Q_task, with a construct graph of dependencies and conflicts and a measurement itemization pipeline Task → Construct → Indicator → Item → Observation → Latent estimate. Any metric (BLEU, CLIPScore, aesthetic model score) is one projection of the space; a construct validity gate (coverage, discriminant validity, convergent evidence, context stability) guards against measuring the wrong thing precisely; the ontology is open — new constructs may be added from residual errors — but every revision is a version, and multimodal quality is not the mean of modality scores.
Definitions
definitions- Core Q_core = (fidelity, coherence, completeness, robustness, usefulness, verifiability); domain schemas for text, image, music, story, design; multimodal coupling dimensions.
- Quality construct graph G_Q = (V_Q, E_Q); construct validity gate G_C; open ontology Q_{t+1} = Q_t ∪ {new construct} under versioning.
- High-ambiguity quality object 𝔔_HA = (Q_schema, G_Q, Q_F, Q_S, θ̂_H, Σ_H, D_R, B_Q, Version_Q).
Assumptions
assumptions- Subjective ≠ unstructured: large populations reliably detect structural failures even in creative artifacts.
Claims
claims- HighAmbiguity ≠ Unmeasurable; Construct ≠ Indicator ≠ Item ≠ Metric; Reliability ≠ Validity.
- Fluency ≠ Factuality; Coherence ≠ Correctness; TechnicalQuality ≠ AestheticQuality; PromptSimilarity ≠ ImageQuality; Novelty ≠ Creativity; AestheticAppeal ≠ Usability.
- MultimodalQuality ≠ Mean(ModalityScores); CrossDomainComparison ⇒ SharedConstructBasis; OntologyRevision ⇒ Versioning; PreciseMeasurement ≠ CorrectConstructSelection.
Formalisation
formalization- Quality fiber view Q = ∪_x Q_x over x = (d, τ, c, a); cross-task projection Π_{x→y} only over shared constructs.
- Creativity as a region (novelty, appropriateness, value, surprise, coherence), not novelty × usefulness.
Predictions
predictions- Benchmarks whose quality ontology drifts without versioning will produce longitudinal comparisons that are silently invalid.
Falsification / failure conditions
falsification_conditions- If a flat, unversioned checklist explains rater residuals and new failure modes as well as the typed construct graph, the ontology machinery is unnecessary.
Known limitations
known_limitations- Schema proposals only; no domain ontology has been validated across populations.
Relations
| Source | Relation | Target | Status | ID |
|---|---|---|---|---|
THY-2026-0108 Typed, versioned quality ontology for high-ambiguity artifacts | extends | THY-2026-0107 Binary residual quality measurement (IBQF / BRQM) | ACTIVE | REL-2026-0373 |
RES-2026-0102 Quality line — measuring outcome quality without asking humans for a score | develops | THY-2026-0108 Typed, versioned quality ontology for high-ambiguity artifacts | ACTIVE | REL-2026-0365 |
THY-2026-0109 Scaffolding capability record: SSR, SDR, SCM and the ablation ladder | extends | THY-2026-0108 Typed, versioned quality ontology for high-ambiguity artifacts | ACTIVE | REL-2026-0376 |
THY-2026-0110 The canonical intelligence event, Pareto comparison and no premature scalarization | extends | THY-2026-0108 Typed, versioned quality ontology for high-ambiguity artifacts | ACTIVE | REL-2026-0379 |
PAP-2026-0108 Paper 08 — How can natural language, images, and creative outputs be measured? A structured quality space for high-ambiguity artifacts | formalizes | THY-2026-0108 Typed, versioned quality ontology for high-ambiguity artifacts | ACTIVE | REL-2026-0420 |
History and provenance
- Canonical URL
- https://evemisslab.com/ai/theory/THY-2026-0108/
- Machine-readable
/ai/theory/THY-2026-0108/index.json- Snapshot
AI-SNAPSHOT-v0.1-fe85b9694a45- Provenance
source- EveMissLab research collection: Intelligence Physical Metrology (真本體論13)
extracted_by- Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports
extracted_at- 2026-09-11
generator- tools/extract_all.py
claim_boundary- status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports