AI Research Laboratory18 records
Results
Observed results, kept apart from their interpretation. Negative, mixed and inconclusive results stay listed.
- ResultRST-2026-0014Real local-model A/B/C table (Qwythos-9B-v2)The v0.2 protocol executed for the first time on a real language model — a local open-weight 9B (Qwythos-9B-v2, Q4_K_M, Ollama, thinking off) as generator, selector and judge; 16 tasks × 2 repetitions × 4 candidates, 380 calls, 32 rows per condition. The PACC runtime (C) gains on supersession alignment (+0.031 vs B, +0.097 vs A) and repair success (+0.070 / +0.094), the two axes the canonical-intent compilation step exists for; derived coherence (+0.011) and intent persistence (−0.002) do not move; valid novelty is slightly lower (−0.045 vs B). Most per-task pairs are ties because the 9B judge saturates near 1.0, and with two repetitions per task the judge's free-text pattern labels never repeat, so creative breadth is not measurable. The conditions chose different candidates in 69 % of task × repetition pairs.
- ResultRST-2026-0015Run 2: no reliable condition difference; breadth not reduced (64 rows)C vs B: supersession +0.0005, repair -0.0125, derived coherence -0.0075, intent -0.0114, valid novelty -0.0187; label-free breadth: cluster entropy C 0.923 / A 0.909 / B 0.864.
- ResultRST-2026-0016Run 3: valid novelty and repair up for the PACC runtime with thinking on; governance flat; breadth not reduced (32 rows)C vs B: valid novelty +0.0687, repair +0.0153, semantic novelty +0.0166, derived coherence -0.0063, intent -0.0053, supersession -0.0059; C vs A: valid novelty +0.0375, repair +0.0466; label-free breadth ratio C 1.047 / A 0.960 / B 0.973.
- ResultRST-2026-0007v0.2 adversarial geometry: N3 converges, N2 crosses the representation gateN3 held-out JS 0.014454 (Level 3); N2 held-out JS 0.053327 > 0.05 (Level 1); N0/N1 0.023921 (Level 3).
- ResultRST-2026-0008v0.4: latent dependence coordinate not recoveredDependence-coordinate pass count 0 for all four systems in every seed; in high correlation real JS ≈ 0.0398 is below the 0.05 absolute gate but no better than shuffled (≈ 0.0400) or constant (≈ 0.0397).
- ResultRST-2026-0009v0.6: composed coordinate 4/4, direct 0/4Composed latent JS ≪ shuffled/constant and ≪ broken composition for every family and geometry (e.g. 0.003603 vs 0.061508 vs 0.058196, moderate N0).
- ResultRST-2026-0010v0.8 basin mapFrozen chart fitted at q* = 0.20: N0/N1/N3 pass on {0.06 … 0.33}, N2 on {0.06 … 0.30}; first failing gate at the edge is negative-control discrimination (ratio ≈ 0.858 at q = 0.36 for N0/N1); local refit 9/9.
- ResultRST-2026-0011v0.10 cocycle: path disagreement ≈ 10⁻⁵ vs control ≈ 10⁻¹Direct-vs-composed latent JS 3.5×10⁻⁵ (N0/N1), 4.0×10⁻⁵ (N2), 2.7×10⁻⁵ (N3); real/shuffled ratio ≈ 3×10⁻⁴; both paths accurate to chart C.
- ResultRST-2026-0012v0.11 reverse destination biasReverse paths agree with each other (latent JS ≈ 10⁻⁵) but both land ≈ 0.011–0.015 from the true chart-A coordinate, above the 0.01 gate, for every family and at every scale tested.
- ResultRST-2026-0013Hybrid v0.1 A/B/C deltasC − B: derived coherence +0.2085, long-horizon retention +0.0142, valid novelty +0.1964, soft-intent −0.0076; C − A pattern entropy −0.4915 (pure creative), recovered to 0.7638 by post-hoc elastic selection.
- ResultRST-2026-0001R1 comparison tableAER-0: 0 stale, 3 recomputations, 0 stable refreshes, 6 workflow reuses, blocks the no-provenance fact, catches the conflicting write. B4: 0 stale, 2 recomputations, 0 stable refreshes, 6 reuses, blocks neither.
- ResultRST-2026-0002R2 classification matrixSeven axes CONVERGED or partially converged (four of them LangGraph richer); six axes AER-default-distinct; graph semantics semantically distinct.
- ResultRST-2026-0003R3: AER-ECT ≈ centralized application gateOn unprovenanced-write blocking, stale-write conflict, automatic audit, dependency binding and single governance site, the centralized application gate and AER-ECT both PASS; the scattered application fails the omission witnesses.
- ResultRST-2026-0004R5 ten-scenario tally5 PREVENTED / 3 OPEN_DETECTED / 1 OPEN_UNDETECTED / 1 SERIALIZED_TO_SEMANTIC_CONFLICT; the undetected case is the coherent full-database forgery.
- ResultRST-2026-0005R6: forgery detection flips under an external anchorThe same coherent full-DB forgery: internal audit OPEN_UNDETECTED, external anchor audit OPEN_DETECTED; prefix truncation: undetected without a trusted head, detected with one.
- ResultRST-2026-0006v0.1 primary table: Level 3 for N0/N1/N2, two familiesHeld-out JS 0.000192 (N0/N1) and 0.008053 (N2) uniform; 0.002878 and 0.010109 heterogeneous; shuffled controls ≈ 0.31–0.33; N0↔N1 agreement 1.000, R² 1.000.
- ResultRST-2026-0101Scaffolding response on three easy tasks: SSR = 1.0, 3.1× device energy at A5, 7.9–9.3× at A2–A4Condition means (quality over available trials / wall s / device J): A0 0.5 / 10.21 / 1715.6; A1 0.5 / 9.87 / 1819.0; A2 0.5 / 79.09 / 15143.0; A3 0.6667 / 99.13 / 15921.5; A4 0.6667 / 69.08 / 13546.2; A5 0.5 / 31.91 / 5321.5. Energy ratio vs A0: A1 1.06, A2 8.827, A3 9.281, A4 7.896, A5 3.102. Total measured GPU energy 0.0891 kWh over 29.9 min of trial time; memory residency rises from 72.1 GiB·s (A0) to 714.6 GiB·s (A3).
- ResultRST-2026-0102Diagnostic: the measured 'quality' was format compliance on 2 of 3 tasks; verifier serialization failed 3/18; the A3/A4 rise is a missingness artifactStrict output-contract compliance 12/36 (MATH-003 12/12, CODE-001 0/12, CON-003 0/12), but the non-canonical post-hoc check finds every completed selected output correct — math 12/12 canonical correct, code 11/11 selected outputs pass all hidden tests after outer Markdown fence removal, constraint 10/10 selected outputs satisfy all constraints after format-only normalization — 33/33 overall. Verifier-enabled trials 18, verifier-protocol failures 3 (A3 2/6, A4 1/6, A5 0/6), each after eight candidates had been generated. Tool calls 0, retries 0; 24 candidates abandoned on abort are neither selected nor discarded in the schema. Quality availability is reported as 22/36 in the aggregate and 33/36 in protocol_compliance.csv. Telemetry sampling wall fraction 0.244.