{
  "section": "results",
  "kind": "result",
  "canonical": "https://evemisslab.com/ai/results/",
  "count": 18,
  "records": [
    {
      "id": "RST-2026-0014",
      "kind": "result",
      "label": "Real local-model A/B/C table (Qwythos-9B-v2)",
      "created_at": "2026-09-11",
      "updated_at": "2026-09-11",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0014/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "The v0.2 protocol executed for the first time on a real language model — a local open-weight 9B (Qwythos-9B-v2, Q4_K_M, Ollama, thinking off) as generator, selector and judge; 16 tasks × 2 repetitions × 4 candidates, 380 calls, 32 rows per condition. The PACC runtime (C) gains on supersession alignment (+0.031 vs B, +0.097 vs A) and repair success (+0.070 / +0.094), the two axes the canonical-intent compilation step exists for; derived coherence (+0.011) and intent persistence (−0.002) do not move; valid novelty is slightly lower (−0.045 vs B). Most per-task pairs are ties because the 9B judge saturates near 1.0, and with two repetitions per task the judge's free-text pattern labels never repeat, so creative breadth is not measurable. The conditions chose different candidates in 69 % of task × repetition pairs.",
        "eml_summary_zh": "v0.2 協定第一次在真實語言模型上執行——本地開放權重 9B（Qwythos-9B-v2、Q4_K_M、Ollama、關閉思考）同時當生成器、選擇器與評審；16 題 × 2 次 × 4 候選，380 次呼叫，每條件 32 筆。PACC runtime（C）在 supersession 對齊（相對 B +0.031、相對 A +0.097）與修復成功率（+0.070／+0.094）上升——正是 canonical intent 編譯步驟存在的那兩個軸；衍生一致性（+0.011）與意圖持續（−0.002）沒有動；有效新穎度略低（相對 B −0.045）。多數逐題配對是平手，因為 9B 評審在接近 1.0 處飽和；每題只重複兩次，評審的自由文字 pattern 標籤從不重複，所以創造廣度量不出來。三個條件在 69 % 的題 × 次配對中選了不同的候選。",
        "eml_label_zh": "真實本地模型 A/B/C 表（Qwythos-9B-v2）",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "REAL MODEL",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "overall": {
            "A_llm_only": {
              "hard_adherence": 0.9875,
              "derived_coherence": 0.9812,
              "intent_persistence": 0.9875,
              "supersession_alignment": 0.8094,
              "repair_success": 0.775,
              "usefulness": 0.9528,
              "semantic_novelty": 0.8084,
              "valid_novelty": 0.8475,
              "literal_check_mean": 0.9792,
              "within_task_pattern_entropy_mean": 1.0
            },
            "B_hard_verifier": {
              "hard_adherence": 0.9859,
              "derived_coherence": 0.9797,
              "intent_persistence": 0.9875,
              "supersession_alignment": 0.875,
              "repair_success": 0.7984,
              "usefulness": 0.9503,
              "semantic_novelty": 0.7725,
              "valid_novelty": 0.8606,
              "literal_check_mean": 0.9792,
              "within_task_pattern_entropy_mean": 1.0
            },
            "C_pacc_runtime": {
              "hard_adherence": 1.0,
              "derived_coherence": 0.9906,
              "intent_persistence": 0.9853,
              "supersession_alignment": 0.9062,
              "repair_success": 0.8688,
              "usefulness": 0.9516,
              "semantic_novelty": 0.7897,
              "valid_novelty": 0.8153,
              "literal_check_mean": 0.9792,
              "within_task_pattern_entropy_mean": 1.0
            }
          },
          "deltas": {
            "C-B": {
              "hard_adherence": 0.0141,
              "derived_coherence": 0.0109,
              "intent_persistence": -0.0022,
              "supersession_alignment": 0.0312,
              "repair_success": 0.0703,
              "usefulness": 0.0012,
              "semantic_novelty": 0.0172,
              "valid_novelty": -0.0453,
              "literal_check_mean": 0.0,
              "within_task_pattern_entropy_mean": 0.0
            },
            "C-A": {
              "hard_adherence": 0.0125,
              "derived_coherence": 0.0094,
              "intent_persistence": -0.0022,
              "supersession_alignment": 0.0969,
              "repair_success": 0.0938,
              "usefulness": -0.0012,
              "semantic_novelty": -0.0188,
              "valid_novelty": -0.0322,
              "literal_check_mean": 0.0,
              "within_task_pattern_entropy_mean": 0.0
            },
            "B-A": {
              "hard_adherence": -0.0016,
              "derived_coherence": -0.0016,
              "intent_persistence": 0.0,
              "supersession_alignment": 0.0656,
              "repair_success": 0.0234,
              "usefulness": -0.0025,
              "semantic_novelty": -0.0359,
              "valid_novelty": 0.0131,
              "literal_check_mean": 0.0,
              "within_task_pattern_entropy_mean": 0.0
            }
          },
          "selection_agreement": {
            "A=B": 0.5625,
            "A=C": 0.40625,
            "B=C": 0.46875,
            "all_same": 0.3125
          }
        },
        "eml_interpretation": "Against the v0.2 predeclared interpretation: the predicted coherence and intent-persistence gains over B are not observed; raw novelty did not decrease (semantic novelty +0.017 vs B); the predicted breadth collapse cannot be tested at this repetition count. What did appear — governance gains on supersession and repair with a valid-novelty cost concentrated in multi_constraint and repair tasks — is mechanism-consistent but small, untested statistically, and runs opposite to the synthetic v0.1 valid-novelty picture (+0.196 there). One model, one run, one same-model judge: a first real data point, not a verdict on the architecture.",
        "eml_limitations": [
          "Descriptive deltas from one local run with a same-model judge."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0014/",
      "json": "/ai/results/RST-2026-0014/index.json"
    },
    {
      "id": "RST-2026-0015",
      "kind": "result",
      "label": "Run 2: no reliable condition difference; breadth not reduced (64 rows)",
      "created_at": "2026-09-11",
      "updated_at": "2026-09-11",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0015/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "C vs B: supersession +0.0005, repair -0.0125, derived coherence -0.0075, intent -0.0114, valid novelty -0.0187; label-free breadth: cluster entropy C 0.923 / A 0.909 / B 0.864.",
        "eml_summary_zh": "C 相對 B：supersession +0.0005、修復 -0.0125、衍生一致性 -0.0075、意圖 -0.0114、有效新穎度 -0.0187；標籤無關廣度：群熵 C 0.923／A 0.909／B 0.864。",
        "eml_label_zh": "第二次執行：條件間無可靠差異；廣度未縮減（64 筆）",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "REAL MODEL",
        "eml_result_type": "NEGATIVE",
        "eml_metrics": {
          "deltas": {
            "C-B": {
              "hard_adherence": -0.0144,
              "derived_coherence": -0.0075,
              "intent_persistence": -0.0114,
              "supersession_alignment": 0.0005,
              "repair_success": -0.0125,
              "usefulness": -0.012,
              "semantic_novelty": 0.0117,
              "valid_novelty": -0.0187,
              "literal_check_mean": -0.0052
            },
            "C-A": {
              "hard_adherence": -0.0222,
              "derived_coherence": -0.0247,
              "intent_persistence": -0.0157,
              "supersession_alignment": -0.0294,
              "repair_success": -0.0423,
              "usefulness": -0.0192,
              "semantic_novelty": -0.0106,
              "valid_novelty": 0.0236,
              "literal_check_mean": 0.0
            },
            "B-A": {
              "hard_adherence": -0.0078,
              "derived_coherence": -0.0172,
              "intent_persistence": -0.0043,
              "supersession_alignment": -0.0298,
              "repair_success": -0.0298,
              "usefulness": -0.0072,
              "semantic_novelty": -0.0223,
              "valid_novelty": 0.0423,
              "literal_check_mean": 0.0052
            }
          },
          "breadth_label_free": {
            "A_llm_only": {
              "cluster_entropy_mean": 0.9091,
              "breadth_ratio_mean": 0.9475,
              "selected_mean_pairwise_distance_mean": 0.1909
            },
            "B_hard_verifier": {
              "cluster_entropy_mean": 0.8635,
              "breadth_ratio_mean": 0.9411,
              "selected_mean_pairwise_distance_mean": 0.1921
            },
            "C_pacc_runtime": {
              "cluster_entropy_mean": 0.9227,
              "breadth_ratio_mean": 0.9745,
              "selected_mean_pairwise_distance_mean": 0.1964
            }
          },
          "selection_agreement": {
            "A=B": 0.546875,
            "A=C": 0.546875,
            "B=C": 0.46875,
            "all_same": 0.34375
          }
        },
        "eml_interpretation": "The first run's governance gains did not replicate; the three conditions are practically equivalent on this model, and the PACC runtime does not narrow creative breadth.",
        "eml_limitations": [
          "Descriptive; same-model judge; one model family; thinking off."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0015/",
      "json": "/ai/results/RST-2026-0015/index.json"
    },
    {
      "id": "RST-2026-0016",
      "kind": "result",
      "label": "Run 3: valid novelty and repair up for the PACC runtime with thinking on; governance flat; breadth not reduced (32 rows)",
      "created_at": "2026-09-11",
      "updated_at": "2026-09-11",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0016/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "C vs B: valid novelty +0.0687, repair +0.0153, semantic novelty +0.0166, derived coherence -0.0063, intent -0.0053, supersession -0.0059; C vs A: valid novelty +0.0375, repair +0.0466; label-free breadth ratio C 1.047 / A 0.960 / B 0.973.",
        "eml_summary_zh": "C 相對 B：有效新穎度 +0.0687、修復 +0.0153、語義新穎度 +0.0166、衍生一致性 -0.0063、意圖 -0.0053、supersession -0.0059；C 相對 A：有效新穎度 +0.0375、修復 +0.0466；標籤無關廣度比 C 1.047／A 0.960／B 0.973。",
        "eml_label_zh": "第三次執行：開啟思考後 PACC runtime 的有效新穎度與修復上升；治理持平；廣度未縮減（32 筆）",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "REAL MODEL",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "deltas": {
            "C-B": {
              "hard_adherence": -0.0069,
              "derived_coherence": -0.0063,
              "intent_persistence": -0.0053,
              "supersession_alignment": -0.0059,
              "repair_success": 0.0153,
              "usefulness": 0.0044,
              "semantic_novelty": 0.0166,
              "valid_novelty": 0.0687,
              "literal_check_mean": 0.0
            },
            "C-A": {
              "hard_adherence": -0.0038,
              "derived_coherence": -0.0013,
              "intent_persistence": -0.0041,
              "supersession_alignment": -0.0034,
              "repair_success": 0.0466,
              "usefulness": -0.0003,
              "semantic_novelty": 0.0128,
              "valid_novelty": 0.0375,
              "literal_check_mean": 0.0
            },
            "B-A": {
              "hard_adherence": 0.0031,
              "derived_coherence": 0.005,
              "intent_persistence": 0.0012,
              "supersession_alignment": 0.0025,
              "repair_success": 0.0312,
              "usefulness": -0.0047,
              "semantic_novelty": -0.0037,
              "valid_novelty": -0.0312,
              "literal_check_mean": 0.0
            }
          },
          "paired_wins_ties_losses": {
            "C-B": {
              "hard_adherence": "1-28-3",
              "derived_coherence": "1-26-5",
              "intent_persistence": "0-29-3",
              "supersession_alignment": "1-27-4",
              "repair_success": "1-29-2",
              "usefulness": "5-22-5",
              "semantic_novelty": "7-18-7",
              "valid_novelty": "6-21-5"
            },
            "C-A": {
              "hard_adherence": "0-29-3",
              "derived_coherence": "2-26-4",
              "intent_persistence": "0-30-2",
              "supersession_alignment": "1-28-3",
              "repair_success": "2-28-2",
              "usefulness": "6-22-4",
              "semantic_novelty": "8-18-6",
              "valid_novelty": "5-20-7"
            }
          },
          "breadth_label_free": {
            "A_llm_only": {
              "cluster_entropy_mean": 0.75,
              "breadth_ratio_mean": 0.9597,
              "selected_mean_pairwise_distance_mean": 0.2079
            },
            "B_hard_verifier": {
              "cluster_entropy_mean": 0.6875,
              "breadth_ratio_mean": 0.9727,
              "selected_mean_pairwise_distance_mean": 0.2087
            },
            "C_pacc_runtime": {
              "cluster_entropy_mean": 0.875,
              "breadth_ratio_mean": 1.0474,
              "selected_mean_pairwise_distance_mean": 0.2275
            }
          },
          "selection_agreement": {
            "A=B": 0.5625,
            "A=C": 0.46875,
            "B=C": 0.46875,
            "all_same": 0.34375
          }
        },
        "eml_interpretation": "The valid-novelty half of the synthetic prediction appears in the means, at a third of the synthetic size, once the model reasons — carried by a few large single-task wins (per-pair 6–21–5 vs B); the coherence/intent half and the breadth-loss prediction do not appear. Unreplicated: a same-size gain in run 1 vanished at 64 rows.",
        "eml_limitations": [
          "Descriptive; 32 rows; same-model judge; one model family; single thinking budget."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0016/",
      "json": "/ai/results/RST-2026-0016/index.json"
    },
    {
      "id": "RST-2026-0007",
      "kind": "result",
      "label": "v0.2 adversarial geometry: N3 converges, N2 crosses the representation gate",
      "created_at": "2026-09-09",
      "updated_at": "2026-09-09",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0007/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "N3 held-out JS 0.014454 (Level 3); N2 held-out JS 0.053327 > 0.05 (Level 1); N0/N1 0.023921 (Level 3).",
        "eml_summary_zh": "N3 held-out JS 0.014454（Level 3）；N2 held-out JS 0.053327 > 0.05（Level 1）；N0/N1 0.023921（Level 3）。",
        "eml_label_zh": "v0.2 對抗幾何：N3 收斂、N2 越過表徵門檻",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "n2_heldout_js": 0.053327,
          "gate": 0.05,
          "n3_heldout_js": 0.014454
        },
        "eml_interpretation": "Convergence is basin-dependent; the preregistered threshold did its job without being moved.",
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0007/",
      "json": "/ai/results/RST-2026-0007/index.json"
    },
    {
      "id": "RST-2026-0008",
      "kind": "result",
      "label": "v0.4: latent dependence coordinate not recovered",
      "created_at": "2026-09-09",
      "updated_at": "2026-09-09",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0008/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Dependence-coordinate pass count 0 for all four systems in every seed; in high correlation real JS ≈ 0.0398 is below the 0.05 absolute gate but no better than shuffled (≈ 0.0400) or constant (≈ 0.0397).",
        "eml_summary_zh": "四個系統在每個 seed 下依賴座標通過數皆為 0；高相關下真實 JS ≈ 0.0398 低於 0.05 絕對門檻，但不比 shuffled（≈ 0.0400）或 constant（≈ 0.0397）好。",
        "eml_label_zh": "v0.4：潛在依賴座標無法重建",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "NEGATIVE",
        "eml_metrics": {
          "dependence_pass": 0,
          "systems": 4
        },
        "eml_interpretation": "Low absolute error ⇏ meaningful coordinate map; the negative-control gate is what stops a near-baseline predictor from being mislabelled latent-state equivalence.",
        "eml_limitations": [
          "'contradicts' the universal direct-affine reading only; the task-level reading survives."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0008/",
      "json": "/ai/results/RST-2026-0008/index.json"
    },
    {
      "id": "RST-2026-0009",
      "kind": "result",
      "label": "v0.6: composed coordinate 4/4, direct 0/4",
      "created_at": "2026-09-09",
      "updated_at": "2026-09-09",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0009/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Composed latent JS ≪ shuffled/constant and ≪ broken composition for every family and geometry (e.g. 0.003603 vs 0.061508 vs 0.058196, moderate N0).",
        "eml_summary_zh": "每個家族與幾何下組合潛在 JS 都 ≪ shuffled／constant 且 ≪ 弄壞的組合（例如中相關 N0：0.003603 vs 0.061508 vs 0.058196）。",
        "eml_label_zh": "v0.6：組合座標 4/4、直接映射 0/4",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "POSITIVE",
        "eml_metrics": {
          "composed_pass": "4/4",
          "direct_pass": "0/4"
        },
        "eml_interpretation": "The latent probability coordinate is hierarchical/compositional relative to these non-probabilistic states rather than directly affine in the raw canonical state.",
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0009/",
      "json": "/ai/results/RST-2026-0009/index.json"
    },
    {
      "id": "RST-2026-0010",
      "kind": "result",
      "label": "v0.8 basin map",
      "created_at": "2026-09-09",
      "updated_at": "2026-09-09",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0010/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Frozen chart fitted at q* = 0.20: N0/N1/N3 pass on {0.06 … 0.33}, N2 on {0.06 … 0.30}; first failing gate at the edge is negative-control discrimination (ratio ≈ 0.858 at q = 0.36 for N0/N1); local refit 9/9.",
        "eml_summary_zh": "在 q* = 0.20 擬合的凍結 chart：N0/N1/N3 在 {0.06 … 0.33} 通過、N2 在 {0.06 … 0.30}；邊緣最先失敗的門檻是負控制判別（N0/N1 在 q = 0.36 比值 ≈ 0.858）；本地重擬合 9/9。",
        "eml_label_zh": "v0.8 吸引域圖",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "N0_N1_N3_first_fail_q": 0.36,
          "N2_first_fail_q": 0.33
        },
        "eml_interpretation": "Shared composed coordinate + family-specific finite basins.",
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0010/",
      "json": "/ai/results/RST-2026-0010/index.json"
    },
    {
      "id": "RST-2026-0011",
      "kind": "result",
      "label": "v0.10 cocycle: path disagreement ≈ 10⁻⁵ vs control ≈ 10⁻¹",
      "created_at": "2026-09-09",
      "updated_at": "2026-09-09",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0011/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Direct-vs-composed latent JS 3.5×10⁻⁵ (N0/N1), 4.0×10⁻⁵ (N2), 2.7×10⁻⁵ (N3); real/shuffled ratio ≈ 3×10⁻⁴; both paths accurate to chart C.",
        "eml_summary_zh": "直接 vs 組合的潛在 JS：3.5×10⁻⁵（N0/N1）、4.0×10⁻⁵（N2）、2.7×10⁻⁵（N3）；真實／shuffled 比 ≈ 3×10⁻⁴；兩條路徑對 chart C 都準確。",
        "eml_label_zh": "v0.10 cocycle：路徑分歧 ≈ 10⁻⁵ vs 控制 ≈ 10⁻¹",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "POSITIVE",
        "eml_metrics": {
          "families_pass": "4/4",
          "secondary_pass_rate": 1.0
        },
        "eml_interpretation": "Forward cocycle coherence is robust where three charts overlap.",
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0011/",
      "json": "/ai/results/RST-2026-0011/index.json"
    },
    {
      "id": "RST-2026-0012",
      "kind": "result",
      "label": "v0.11 reverse destination bias",
      "created_at": "2026-09-09",
      "updated_at": "2026-09-09",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0012/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Reverse paths agree with each other (latent JS ≈ 10⁻⁵) but both land ≈ 0.011–0.015 from the true chart-A coordinate, above the 0.01 gate, for every family and at every scale tested.",
        "eml_summary_zh": "反向路徑彼此一致（潛在 JS ≈ 10⁻⁵），但兩者都落在離真正 chart-A 座標 ≈ 0.011–0.015 處，高於 0.01 門檻，所有家族、所有測試規模皆然。",
        "eml_label_zh": "v0.11 反向目的偏差",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "reverse_pass": "0/4 primary+secondary, 0/3 large",
          "path_coherence": "≈1e-5"
        },
        "eml_interpretation": "Path coherence is not enough for atlas closure; the bias is the object v0.12–v0.13 then tried to explain.",
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0012/",
      "json": "/ai/results/RST-2026-0012/index.json"
    },
    {
      "id": "RST-2026-0013",
      "kind": "result",
      "label": "Hybrid v0.1 A/B/C deltas",
      "created_at": "2026-09-09",
      "updated_at": "2026-09-09",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0013/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "C − B: derived coherence +0.2085, long-horizon retention +0.0142, valid novelty +0.1964, soft-intent −0.0076; C − A pattern entropy −0.4915 (pure creative), recovered to 0.7638 by post-hoc elastic selection.",
        "eml_summary_zh": "C − B：衍生一致性 +0.2085、長程保持 +0.0142、有效新穎度 +0.1964、軟意圖 −0.0076；C − A pattern entropy −0.4915（純創意），事後 elastic 選擇恢復到 0.7638。",
        "eml_label_zh": "Hybrid v0.1 A/B/C 差值",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "derived_delta_C_B": 0.2085,
          "valid_novelty_delta_C_B": 0.1964,
          "soft_delta_C_B": -0.0076,
          "entropy_delta_C_A": -0.4915,
          "elastic_entropy": 0.7638
        },
        "eml_interpretation": "Gains come from derived dependency coherence and constraint-satisfying novelty, not from more rejection; the breadth cost is mitigable without relaxing commit constraints.",
        "eml_limitations": [
          "Synthetic scoring; a real-model result is required before any claim about language models."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0013/",
      "json": "/ai/results/RST-2026-0013/index.json"
    },
    {
      "id": "RST-2026-0001",
      "kind": "result",
      "label": "R1 comparison table",
      "created_at": "2026-09-08",
      "updated_at": "2026-09-08",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E2",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0001/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "AER-0: 0 stale, 3 recomputations, 0 stable refreshes, 6 workflow reuses, blocks the no-provenance fact, catches the conflicting write. B4: 0 stale, 2 recomputations, 0 stable refreshes, 6 reuses, blocks neither.",
        "eml_summary_zh": "AER-0：0 過期、3 次重算、0 次穩定節點刷新、6 次 workflow 重用、擋下無 provenance 事實、抓到衝突寫入。B4：0 過期、2 次重算、0 次穩定刷新、6 次重用、兩者都不擋。",
        "eml_label_zh": "R1 比較表",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "DETERMINISTIC RUNTIME",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "aer_recomputations": 3,
          "b4_recomputations": 2,
          "aer_only": [
            "blocks no-provenance fact",
            "VersionConflict on stale expected version"
          ]
        },
        "eml_interpretation": "Tension scheduling paid one proactive refresh that pure event invalidation did not need in an event-rich world; the surviving distinction is mutation authority, not refresh efficiency.",
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0001/",
      "json": "/ai/results/RST-2026-0001/index.json"
    },
    {
      "id": "RST-2026-0002",
      "kind": "result",
      "label": "R2 classification matrix",
      "created_at": "2026-09-08",
      "updated_at": "2026-09-08",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E2",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0002/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Seven axes CONVERGED or partially converged (four of them LangGraph richer); six axes AER-default-distinct; graph semantics semantically distinct.",
        "eml_summary_zh": "七個軸收斂或部分收斂（其中四個 LangGraph 更豐富）；六個軸為 AER 預設獨有；圖語義為語義上不同。",
        "eml_label_zh": "R2 分類矩陣",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "DETERMINISTIC RUNTIME",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "converged_or_partial": 7,
          "langgraph_richer": 4,
          "aer_default_distinct": 6
        },
        "eml_interpretation": "Supports the attractor hypothesis for external state, persistence, graph/workflow execution, history and durable recovery; moves the remaining disagreement upward into what state means, who may mutate it and what a mutation requires.",
        "eml_limitations": [
          "Documentary evidence on the LangGraph side; no executable replay."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0002/",
      "json": "/ai/results/RST-2026-0002/index.json"
    },
    {
      "id": "RST-2026-0003",
      "kind": "result",
      "label": "R3: AER-ECT ≈ centralized application gate",
      "created_at": "2026-09-08",
      "updated_at": "2026-09-08",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E2",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0003/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "On unprovenanced-write blocking, stale-write conflict, automatic audit, dependency binding and single governance site, the centralized application gate and AER-ECT both PASS; the scattered application fails the omission witnesses.",
        "eml_summary_zh": "在擋無 provenance 寫入、過期寫入衝突、自動 audit、依賴綁定與單一治理站點上，集中式應用 gate 與 AER-ECT 都 PASS；分散式應用在遺漏見證上失敗。",
        "eml_label_zh": "R3：AER-ECT ≈ 集中式應用 gate",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "DETERMINISTIC RUNTIME",
        "eml_result_type": "NEGATIVE",
        "eml_metrics": {
          "aer_ect_vs_central_gate": "equal on all tested semantics",
          "scattered_failures": 3
        },
        "eml_interpretation": "The most important negative result of the line: what AER-ECT computes can be reproduced by an application; what changes is where the obligation lives.",
        "eml_limitations": [
          "'contradicts' here means the computational-uniqueness reading of ECT; the architectural-elevation reading survives."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0003/",
      "json": "/ai/results/RST-2026-0003/index.json"
    },
    {
      "id": "RST-2026-0004",
      "kind": "result",
      "label": "R5 ten-scenario tally",
      "created_at": "2026-09-08",
      "updated_at": "2026-09-08",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E2",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0004/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "5 PREVENTED / 3 OPEN_DETECTED / 1 OPEN_UNDETECTED / 1 SERIALIZED_TO_SEMANTIC_CONFLICT; the undetected case is the coherent full-database forgery.",
        "eml_summary_zh": "5 PREVENTED／3 OPEN_DETECTED／1 OPEN_UNDETECTED／1 SERIALIZED_TO_SEMANTIC_CONFLICT；未偵測的那一例是連貫的整庫偽造。",
        "eml_label_zh": "R5 十情境統計",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "DETERMINISTIC RUNTIME",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "prevented": 5,
          "open_detected": 3,
          "open_undetected": 1,
          "serialized": 1
        },
        "eml_interpretation": "No external trust anchor ⇒ no full-DB tamperproof claim — a structural result that defines R6.",
        "eml_limitations": [
          "Retained negative witnesses are part of the verdict, not implementation accidents."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0004/",
      "json": "/ai/results/RST-2026-0004/index.json"
    },
    {
      "id": "RST-2026-0005",
      "kind": "result",
      "label": "R6: forgery detection flips under an external anchor",
      "created_at": "2026-09-08",
      "updated_at": "2026-09-08",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E2",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0005/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "The same coherent full-DB forgery: internal audit OPEN_UNDETECTED, external anchor audit OPEN_DETECTED; prefix truncation: undetected without a trusted head, detected with one.",
        "eml_summary_zh": "同一個連貫整庫偽造：內部稽核 OPEN_UNDETECTED、外部 anchor 稽核 OPEN_DETECTED；前綴截斷：無可信 head 未偵測、有則偵測。",
        "eml_label_zh": "R6：外部 anchor 下偽造偵測翻轉",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "DETERMINISTIC RUNTIME",
        "eml_result_type": "POSITIVE",
        "eml_metrics": {
          "forgery_internal": "OPEN_UNDETECTED",
          "forgery_external": "OPEN_DETECTED",
          "truncation_without_head": "OPEN_UNDETECTED",
          "truncation_with_head": "OPEN_DETECTED"
        },
        "eml_interpretation": "Detection, not prevention; and only under the stated key and head trust assumptions.",
        "eml_limitations": [
          "Not a Certificate Transparency or Rekor implementation; a smaller signed hash chain with an optional head receipt."
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0005/",
      "json": "/ai/results/RST-2026-0005/index.json"
    },
    {
      "id": "RST-2026-0006",
      "kind": "result",
      "label": "v0.1 primary table: Level 3 for N0/N1/N2, two families",
      "created_at": "2026-09-08",
      "updated_at": "2026-09-08",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E3",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0006/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_aes/extract.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Held-out JS 0.000192 (N0/N1) and 0.008053 (N2) uniform; 0.002878 and 0.010109 heterogeneous; shuffled controls ≈ 0.31–0.33; N0↔N1 agreement 1.000, R² 1.000.",
        "eml_summary_zh": "均勻：held-out JS 0.000192（N0/N1）與 0.008053（N2）；異質：0.002878 與 0.010109；shuffled 控制組 ≈ 0.31–0.33；N0↔N1 一致度 1.000、R² 1.000。",
        "eml_label_zh": "v0.1 主要表：N0/N1/N2 皆 Level 3，兩個家族",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0001",
        "eml_data_basis": "SYNTHETIC",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "level": {
            "N0": 3,
            "N1": 3,
            "N2": 3
          },
          "families": [
            [
              "N0",
              "N1"
            ],
            [
              "N2"
            ]
          ]
        },
        "eml_interpretation": "PACC-B/R/D supported as a controlled micro-environment witness; PACC-A not yet supported.",
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ],
        "eml_ai_collaborators": [
          "Sol (GPT-5.6, OpenAI ChatGPT)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0006/",
      "json": "/ai/results/RST-2026-0006/index.json"
    },
    {
      "id": "RST-2026-0101",
      "kind": "result",
      "label": "Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4",
      "created_at": "2026-09-03",
      "updated_at": "2026-09-07",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E2",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0101/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Intelligence Physical Metrology (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_all.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Condition means (quality over available trials / wall s / device J): A0 0.5 / 10.21 / 1715.6; A1 0.5 / 9.87 / 1819.0; A2 0.5 / 79.09 / 15143.0; A3 0.6667 / 99.13 / 15921.5; A4 0.6667 / 69.08 / 13546.2; A5 0.5 / 31.91 / 5321.5. Energy ratio vs A0: A1 1.06, A2 8.827, A3 9.281, A4 7.896, A5 3.102. Total measured GPU energy 0.0891 kWh over 29.9 min of trial time; memory residency rises from 72.1 GiB·s (A0) to 714.6 GiB·s (A3).",
        "eml_summary_zh": "各條件平均（可用試驗的品質／wall 秒／裝置焦耳）：A0 0.5／10.21／1715.6；A1 0.5／9.87／1819.0；A2 0.5／79.09／15143.0；A3 0.6667／99.13／15921.5；A4 0.6667／69.08／13546.2；A5 0.5／31.91／5321.5。能量相對 A0：A1 1.06、A2 8.827、A3 9.281、A4 7.896、A5 3.102。36 次試驗共量得 GPU 能量 0.0891 kWh、試驗時間 29.9 分鐘；記憶體駐留從 A0 的 72.1 GiB·s 升到 A3 的 714.6 GiB·s。",
        "eml_label_zh": "三個簡單任務上的鷹架響應：SSR = 1.0，A5 的裝置能量 3.1×、A2–A4 7.9–9.3×",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0101",
        "eml_data_basis": "REAL MODEL",
        "eml_result_type": "MIXED",
        "eml_metrics": {
          "ssr": 1.0,
          "sdr": 0.0,
          "scm": {
            "device_energy_j": 3.1019,
            "wall_time_s": 3.1264
          },
          "by_condition": {
            "A0": {
              "quality_mean": 0.5,
              "quality_n": 4,
              "success_rate": 0.5,
              "wall_time_s": 10.21,
              "device_energy_j": 1715.6,
              "energy_ratio_vs_A0": 1.0,
              "gpu_peak_memory_gib": 7.23,
              "gpu_memory_residency_gib_s": 72.1,
              "gpu_utilization_integral_s": 6.57
            },
            "A1": {
              "quality_mean": 0.5,
              "quality_n": 4,
              "success_rate": 0.5,
              "wall_time_s": 9.87,
              "device_energy_j": 1819.0,
              "energy_ratio_vs_A0": 1.06,
              "gpu_peak_memory_gib": 7.32,
              "gpu_memory_residency_gib_s": 70.0,
              "gpu_utilization_integral_s": 6.26
            },
            "A2": {
              "quality_mean": 0.5,
              "quality_n": 4,
              "success_rate": 0.5,
              "wall_time_s": 79.09,
              "device_energy_j": 15143.0,
              "energy_ratio_vs_A0": 8.827,
              "gpu_peak_memory_gib": 7.28,
              "gpu_memory_residency_gib_s": 570.0,
              "gpu_utilization_integral_s": 54.09
            },
            "A3": {
              "quality_mean": 0.6667,
              "quality_n": 3,
              "success_rate": 0.6667,
              "wall_time_s": 99.13,
              "device_energy_j": 15921.5,
              "energy_ratio_vs_A0": 9.281,
              "gpu_peak_memory_gib": 7.34,
              "gpu_memory_residency_gib_s": 714.6,
              "gpu_utilization_integral_s": 70.02
            },
            "A4": {
              "quality_mean": 0.6667,
              "quality_n": 3,
              "success_rate": 0.6667,
              "wall_time_s": 69.08,
              "device_energy_j": 13546.2,
              "energy_ratio_vs_A0": 7.896,
              "gpu_peak_memory_gib": 7.28,
              "gpu_memory_residency_gib_s": 496.9,
              "gpu_utilization_integral_s": 46.93
            },
            "A5": {
              "quality_mean": 0.5,
              "quality_n": 4,
              "success_rate": 0.5,
              "wall_time_s": 31.91,
              "device_energy_j": 5321.5,
              "energy_ratio_vs_A0": 3.102,
              "gpu_peak_memory_gib": 7.24,
              "gpu_memory_residency_gib_s": 228.7,
              "gpu_utilization_integral_s": 22.17
            }
          },
          "marginal_yield": {
            "A0->A1": {
              "allocated_device_time_s": null,
              "device_energy_j": 0.0,
              "wall_time_s": null
            },
            "A1->A2": {
              "allocated_device_time_s": null,
              "device_energy_j": 0.0,
              "wall_time_s": 0.0
            },
            "A2->A3": {
              "allocated_device_time_s": null,
              "device_energy_j": 0.00021406648843853404,
              "wall_time_s": 0.008315919866432606
            },
            "A3->A4": {
              "allocated_device_time_s": null,
              "device_energy_j": null,
              "wall_time_s": null
            },
            "A4->A5": {
              "allocated_device_time_s": null,
              "device_energy_j": null,
              "wall_time_s": null
            }
          }
        },
        "eml_interpretation": "The cost side of the scaffolding response curve is real and steep; the quality side is flat because the tasks were already solved at A0 and because the quality axis was confounded (RST-2026-0102). This is one point on the 'no gap' side of F4 with almost no weight: it neither supports nor refutes the scaffolding-separation hypothesis on non-trivial tasks. The A3/A4 0.667 is not a gain: the aborted CON-003 trials dropped out of the quality denominator and the surviving mean rose.",
        "eml_limitations": [
          "quality_n is 4 (A0–A2, A5) or 3 (A3, A4) per condition because CODE-001 quality is unavailable and three trials aborted — means over 3–4 values.",
          "Device-measured GPU energy at ~24 % sampling overhead; not marginal energy."
        ],
        "eml_ai_collaborators": [
          "Aletheia (GPT-5.6 Sol, OpenAI ChatGPT) — 2026-09-07 diagnostic",
          "Splice (Claude Code, Anthropic) — execution and RESULT note"
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0101/",
      "json": "/ai/results/RST-2026-0101/index.json"
    },
    {
      "id": "RST-2026-0102",
      "kind": "result",
      "label": "Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifact",
      "created_at": "2026-09-03",
      "updated_at": "2026-09-07",
      "values": {
        "eml_status": "STABLE",
        "eml_evidence_level": "E2",
        "eml_object_version": "0.1",
        "eml_canonical_url": "https://evemisslab.com/ai/results/RST-2026-0102/",
        "eml_provenance": {
          "source": "EveMissLab research collection: Intelligence Physical Metrology (真本體論13)",
          "extracted_by": "Splice (Claude Code), reading the canonical UTF-8 sources and each package's own reports",
          "extracted_at": "2026-09-11",
          "generator": "tools/extract_all.py",
          "claim_boundary": "status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports"
        },
        "eml_summary": "Strict output-contract compliance 12/36 (MATH-003 12/12, CODE-001 0/12, CON-003 0/12), but the non-canonical post-hoc check finds every completed selected output correct — math 12/12 canonical correct, code 11/11 selected outputs pass all hidden tests after outer Markdown fence removal, constraint 10/10 selected outputs satisfy all constraints after format-only normalization — 33/33 overall. Verifier-enabled trials 18, verifier-protocol failures 3 (A3 2/6, A4 1/6, A5 0/6), each after eight candidates had been generated. Tool calls 0, retries 0; 24 candidates abandoned on abort are neither selected nor discarded in the schema. Quality availability is reported as 22/36 in the aggregate and 33/36 in protocol_compliance.csv. Telemetry sampling wall fraction 0.244.",
        "eml_summary_zh": "嚴格輸出契約合規 12/36（MATH-003 12/12、CODE-001 0/12、CON-003 0/12），但非 canonical 的事後檢查發現所有完成的選中輸出都正確——數學 12/12 canonical correct、程式 11/11 selected outputs pass all hidden tests after outer Markdown fence removal、約束 10/10 selected outputs satisfy all constraints after format-only normalization——整體 33/33。啟用驗證器的試驗 18 次，驗證器協定失敗 3 次（A3 2/6、A4 1/6、A5 0/6），每次都在已生成八個候選之後。工具呼叫 0 次、重試 0 次；24 個在中止時被拋棄的候選在 schema 裡既非選中也非丟棄。品質可得率在聚合檔報 22/36、在 protocol_compliance.csv 報 33/36。遙測取樣占 wall time 的 0.244。",
        "eml_label_zh": "診斷：三題中兩題量到的「品質」是格式服從性；驗證器序列化失敗 3/18；A3/A4 的上升是缺值假象",
        "eml_primary_domain": "Evaluation",
        "eml_program_id": "PRG-2026-0101",
        "eml_data_basis": "REAL MODEL",
        "eml_result_type": "NEGATIVE",
        "eml_metrics": {
          "strict_protocol_compliance": {
            "count": 12,
            "total": 36,
            "rate": 0.3333333333333333
          },
          "posthoc_semantic_diagnostic": {
            "canonical": false,
            "math": "12/12 canonical correct",
            "code": "11/11 selected outputs pass all hidden tests after outer Markdown fence removal",
            "constraint": "10/10 selected outputs satisfy all constraints after format-only normalization",
            "completed_selected_outputs_correct": "33/33"
          },
          "verifier_failure": {
            "count": 3,
            "verifier_enabled_trials": 18,
            "rate": 0.16666666666666666,
            "by_condition": {
              "A3": "2/6",
              "A4": "1/6",
              "A5": "0/6"
            }
          },
          "operational_totals": {
            "model_invocations": 186,
            "trajectories": 162,
            "retries": 0,
            "tool_calls": 0,
            "verifier_passes": 18,
            "candidates_created": 162,
            "candidates_selected": 33,
            "candidates_discarded": 105,
            "failure_events": 6,
            "candidates_abandoned_on_abort": 24
          },
          "quality_availability_reporting": {
            "aggregate_analysis": "22/36",
            "protocol_compliance_csv": "33/36",
            "inconsistency": true
          },
          "instrumentation": {
            "mean_sampling_call_wall_fraction": 0.2436,
            "target_sampling_ms": 250,
            "observed_cadence_ms": "335–350"
          },
          "instrument_revisions_required_before_XA-07": [
            "R1 typed subject/system failure is valid data, separate from instrument failure",
            "R2 split task_semantic_quality / output_contract_compliance / system_completion_reliability / verifier_protocol_reliability",
            "R3 failure-aware aggregation (no survivor means)",
            "R4 abandoned-candidate accounting",
            "R5 preserve verifier parse diagnostics",
            "R6 cheaper telemetry (persistent nvidia-smi / NVML)",
            "R7 tool-trigger tasks"
          ]
        },
        "eml_interpretation": "Negative for the instrument, informative for the theory. Task semantic quality ≠ protocol/serialization compliance — the model solved everything it completed and was scored 0.5 on average for not obeying a JSON/no-fence contract; same-model verification reduced system reliability rather than raising quality; tool access ≠ tool utilization (enabled, never used). The bundle stays immutable as XA-06 empirical v0.1; the next step is to revise the instrument and re-run a small validation set, not to re-roll failures until they disappear.",
        "eml_limitations": [
          "The post-hoc semantic check is explicitly non-canonical and does not make the strict outputs compliant.",
          "One model; whether stronger models obey the output contract is unknown."
        ],
        "eml_ai_collaborators": [
          "Aletheia (GPT-5.6 Sol, OpenAI ChatGPT) — 2026-09-07 diagnostic",
          "Splice (Claude Code, Anthropic) — execution and RESULT note"
        ],
        "eml_authors": [
          "Neo.K (EveMissLab)"
        ]
      },
      "canonical_url": "https://evemisslab.com/ai/results/RST-2026-0102/",
      "json": "/ai/results/RST-2026-0102/index.json"
    }
  ],
  "snapshot": {
    "snapshot_id": "AI-SNAPSHOT-v0.1-fe85b9694a45",
    "created_at": "2026-09-11T05:00:27Z",
    "format_version": "0.1",
    "sedb_baseline": "v0.4B contract; static source content/ai/",
    "generator_version": "evemisslab-com ai_research 0.1",
    "object_count": 124,
    "relation_count": 499,
    "artifact_count": 58
  }
}
