AI Research Laboratory32 records
Experiments
Every experiment is a first-class record: setup, procedure, runs, observed results, and an interpretation kept separate from them.
- ExperimentEXP-2026-0023PACC-Hybrid v0.2 — first real-model run, on a local 9B open-weight modelThe v0.2 protocol executed for the first time on a real language model — a local open-weight 9B (Qwythos-9B-v2, Q4_K_M, Ollama, thinking off) as generator, selector and judge; 16 tasks × 2 repetitions × 4 candidates, 380 calls, 32 rows per condition. The PACC runtime (C) gains on supersession alignment (+0.031 vs B, +0.097 vs A) and repair success (+0.070 / +0.094), the two axes the canonical-intent compilation step exists for; derived coherence (+0.011) and intent persistence (−0.002) do not move; valid novelty is slightly lower (−0.045 vs B). Most per-task pairs are ties because the 9B judge saturates near 1.0, and with two repetitions per task the judge's free-text pattern labels never repeat, so creative breadth is not measurable. The conditions chose different candidates in 69 % of task × repetition pairs.
- ExperimentEXP-2026-0024PACC-Hybrid v0.2 — real-model run 2: four repetitions and label-free creative breadthSame model and settings as the first run, repetitions raised from two to four: 16 tasks × 4 × 4 candidates, 754 unique model requests, 64 rows per condition. The PACC runtime (C) shows no reliable advantage on any judge axis — supersession +0.0005 and repair -0.0125 vs B, so the first run's gains did not replicate — and sits a few hundredths below A and B on adherence, coherence and intent persistence. A new label-free breadth measure (local nomic-embed-text embeddings, k-means labels over each task's candidate pool) finds no collapse: C's cluster entropy 0.923 vs A 0.909 / B 0.864, breadth ratio 0.975 vs 0.948 / 0.941. Five selector/judge outputs failed strict JSON parsing and were regenerated under a disclosed retry policy.
- ExperimentEXP-2026-0025PACC-Hybrid v0.2 — real-model run 3: thinking enabled with enlarged output budgetsThe same local 9B model with its thinking enabled (reasoning.effort = medium) and 3,500 extra output tokens on every call so the hidden reasoning has room; 12k context; 16 tasks × 2 repetitions × 4 candidates, 380 unique requests, 32 rows per condition, 3.7 h. The PACC runtime (C) has the highest mean valid novelty (0.8375; +0.0687 vs B, +0.0375 vs A) and mean repair success (+0.0153 / +0.0466) — but per task × repetition the record is even (valid novelty wins–ties–losses 6-21-5 vs B, 5-20-7 vs A; repair 1-29-2 vs B), so the means come from a few large single-task differences in multi_constraint and repair tasks — while adherence, coherence, intent persistence and supersession are flat to a few thousandths lower (-0.0069, -0.0063, -0.0053, -0.0059 vs B). All three conditions pass every deterministic literal check with thinking on. Label-free breadth is again not reduced under C: cluster entropy 0.875 vs A 0.750 / B 0.688, breadth ratio 1.047 vs 0.960 / 0.973. Two selector outputs were regenerated under the disclosed retry policy (one cut off, one empty after the reasoning consumed its whole 3,800-token budget).
- ExperimentEXP-2026-0009PACC-Lab v0.2 — third family (N3) and adversarial evidence geometryAdds N3 constraint competition, a nonlinear constraint-field family, plus an adversarial geometry chosen to break additive/log-odds mappings, keeping every v0.1 threshold. N3 reaches Level 3 in all three geometries; N2 keeps agreement 0.8546 and small update/intervention error but its held-out representation JS 0.053327 crosses the 0.05 gate under adversarial geometry, so it is classified behavioural convergence only. Independent families under the linear diagnostic: {N0, N1}, {N2}, {N3}; robust across all three geometries: two.
- ExperimentEXP-2026-0010PACC-Lab v0.3 — learned source reliability without an oracleBoth sides lose the source-quality oracle: the Bayesian reference learns Beta-Bernoulli source quality; the non-probabilistic systems learn a qualitative reputation (trust, friction, streak, familiarity). Fitted on training worlds and evaluated on unseen source-quality permutations, the reputation state maps to the Beta-Bernoulli state with JS ≈ 0.00714 versus 0.017–0.018 for shuffled/constant controls, and feedback-update commutation ≈ 0.00130 — stable 6/6 across seeds. Task-state convergence is architecture- and seed-sensitive: at exact primary scale N0/N1 4/4, N2 3/4, N3 2/4.
- ExperimentEXP-2026-0011PACC-Lab v0.4 — correlated sources and dependence geometrySources are grouped into clusters with a latent shared inversion; the correct reference uses the joint likelihood and a naive independent-product Bayes is the negative control (joint beats naive in moderate and high geometry; gap exactly 0 in the independent geometry). Non-probabilistic systems see only cluster membership. Task-state convergence survives: N0/N1/N3 pass in moderate and high correlation (N2 drops to agreement 0.8364 in high). The latent dependence coordinate — the posterior that the cluster is in its corrupted branch — is not recovered by a train-only affine-sigmoid map from any system: real JS ≈ shuffled ≈ constant.
- ExperimentEXP-2026-0012PACC-Lab v0.5 — does a richer relation state rescue the latent coordinate?A rich relation graph keeps far more source/pair/pattern information than the compact v0.4 bundle while leaving task actions identical. It does not rescue the frozen direct affine-sigmoid map to the common-cause coordinate: compact 0 and rich 0 robust passes, real/control ratios ≈ 0.96–1.02, and no system rescued in four secondary seeds. Counterfactual source-flip probes often pass alone, showing partial local directional alignment without a globally discriminative coordinate.
- ExperimentEXP-2026-0013PACC-Lab v0.6 — hierarchical composed coordinateA preregistered 95-dimensional composed coordinate — the validated task-state map's held-out decision quotient combined with the current relation pattern and a bounded set of interaction terms — predicts the Bayesian common-cause coordinate for all four families in both geometries: composed latent JS 0.0034–0.0073 versus 0.040–0.062 for shuffled/constant and 0.040–0.060 for a deliberately broken composition; counterfactual JS 0.0018–0.0047; four secondary seeds give direct 0.0 and composed 1.0 pass rates.
- ExperimentEXP-2026-0014PACC-Lab v0.7 — cross-geometry coordinate transfer without refittingThe v0.6 composed coordinate is frozen on one dependence geometry and applied to the other. N0/N1/N3 transfer bidirectionally on the primary protocol and 4/4 secondary seeds; N2 passes high → moderate but misses the moderate → high task-quotient gate by ≈ 0.0003 on the primary seed (2/4 secondary), while a targeted larger stress passes 3/3. Every target-local refit passes, so N2's failure is not a target-chart failure.
- ExperimentEXP-2026-0015PACC-Lab v0.8 — frozen-coordinate transfer basins over a correlation sweepOne composed coordinate fitted at q* = 0.20 is applied unchanged across a nine-point sweep q ∈ {0.06 … 0.42} (within-cluster correctness correlation rising from ≈ 0.203 to ≈ 0.502). N0/N1/N3 pass through q = 0.33 and fail from 0.36; N2 passes through 0.30 and fails from 0.33. At the edge the negative-control discrimination margin is lost first while absolute latent error stays small; target-local refit passes 9/9 everywhere. Four secondary seeds reproduce the basin exactly for N0/N1/N3 and a seed-sensitive N2 edge at 0.33.
- ExperimentEXP-2026-0016PACC-Lab v0.9 — a three-anchor probability-coordinate atlasThree preregistered charts at q = 0.12, 0.24, 0.36 give every family the same union {0.06 … 0.36}: 8/9 sweep coverage, stable across three secondary seeds, never covering q = 0.42. Nearest charts (0.12/0.24) agree directly; distant charts (0.12/0.36) do not (task JS ≈ 0.071–0.082), yet low-complexity transition maps calibrated on an independent seed pass 5/6 directions per family on the primary protocol — the weak direction is 0.36 → 0.12 (latent JS ≈ 0.017–0.021 > 0.01). Transition coherence is not robust across secondary seeds (mean pass ≈ 0.53–0.67).
- ExperimentEXP-2026-0017PACC-Lab v0.10 — oriented cocycle coherence on frozen triple overlapOn the frozen region where charts A, B and C were all valid in the v0.9 atlas, the direct transition A → C and the composed A → B → C agree to ≈ 10⁻⁵ in latent JS (shuffled-pair control ≈ 10⁻¹), while both paths separately stay accurate to chart C (latent JS ≈ 0.0045–0.0054). All four families pass on the primary seed and on four secondary seeds. N2 has only one triple-overlap geometry (q = 0.24), so its witness is narrower.
- ExperimentEXP-2026-0018PACC-Lab v0.11 — bidirectional cocycle and inverse consistencySeparates three properties: forward cocycle composition (survives, all families), round-trip inverse consistency (all pairwise round trips pass on the primary run, recover 3/3 at larger scale, but are sample-sensitive at small scale, weakest on the widest pair AC), and reverse transport fidelity to the destination chart (fails for every family: direct C → A latent fidelity 0.0116–0.0148 > 0.01 while direct and composed reverse paths agree to ≈ 10⁻⁵).
- ExperimentEXP-2026-0019PACC-Lab v0.12 — bounded quadratic transitionTransition capacity rises in one preregistered step from 20 affine to 60 fixed degree-2 coefficients with charts, support, split, controls and thresholds unchanged. No family is rescued: reverse fidelity 0.0111 (N0/N1), 0.0139 (N2), 0.0141 (N3) versus the 0.01 gate, 0/4 secondary and 0/3 at larger scale; forward cocycle and inverse consistency survive the quadratic class. The lab stops here rather than escalating to cubic, quartic or neural transitions after seeing the result.
- ExperimentEXP-2026-0020PACC-Lab v0.13 — reverse residual field / base-point dependenceThe held-out reverse residual r_CA(S) = Φ_A(S) − T_CA(Φ_C(S)) is measured for mean, covariance, latent-logit energy concentration, between- versus within-stratum variance, direction stability and a train-only per-stratum constant correction against global and shuffled-stratum controls. N0/N1 concentrate 98 % of residual energy on the latent-logit axis, yet the mean direction flips sign across seeds (cross-seed cosine ≈ −1), between-q structure is ≈ 0.1–0.3 % against a 20 % gate, and stratum correction changes held-out fidelity by < 10⁻⁵ — no family passes.
- ExperimentEXP-2026-0021PACC-Hybrid v0.1 — synthetic A/B/C architecture witnessOver 1,568 identical candidate pools, generator-only (A), hard verifier (B) and PACC runtime (C) are scored on hard adherence, derived coherence, soft-intent satisfaction, long-horizon retention, novelty and pattern entropy. B already saturates literal hard adherence (1.0); C adds derived coherence +0.2085, long-horizon retention +0.0142 and valid novelty +0.1964 over B, loses 0.0076 soft-intent satisfaction, and in the pure-creative control raises pointwise novelty (0.9864 vs 0.8889) while collapsing pattern entropy (0.2284 vs 0.7199). A post-hoc 'elastic' exploration policy (D) restores entropy to 0.7638 with derived coherence still 1.0. Zero real LLM calls; eight seeds sign-stable.
- ExperimentEXP-2026-0022PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)Sixteen hand-authored natural-language tasks across the preregistered families; A/B/C share one candidate ledger and equal accounted budgets; C receives no gold constraint or supersession metadata; the gold rubric is visible only to a condition-blind judge; literal machine checks are independent of the judge; identical answers reuse one judge cache key; the system fails closed without OPENAI_API_KEY; a frozen-cache replay provider allows exact replay after a live run. 25 tests pass. The execution runtime had no API key, so no real model output exists; the bundled mock smoke file is marked MOCK_ONLY_NOT_REAL_MODEL.
- ExperimentEXP-2026-0001AER-0 MVP v0.1 closure: are the invariants executable?The approved Python + SQLite runtime was built under mssp-tdd-apr and closed on behavioural, structural and discriminative witnesses: candidate gate, provenance validation, stale-version conflict, stable-vs-volatile scheduling, shock activation, verifier rejection, workflow reuse/adapt/create, model swap, model-free core, end-to-end verified commit. 25 tests; the refresh benchmark selected 2 nodes under fixed TTL versus 1 volatile node under tension scheduling; a clean extracted replay of the sealed archive passed before release.
- ExperimentEXP-2026-0002R1 — deterministic semantics comparison against four baselinesTwo facts (one stable, one changing at hours 8 and 16), eight queries over 17 simulated hours. B1 stateless recomputes every query; B2 fixed 6-hour TTL; B3 adds workflow memory; B4 an evented compound agent with event-driven invalidation. AER-0 matches B4 on zero stale answers, zero stable refreshes and workflow reuse, needs one more recomputation (3 vs 2) because its tension policy refreshes proactively, and is alone in blocking a no-provenance fact and catching a conflicting write.
- ExperimentEXP-2026-0003R2 — source-grounded structural comparison with LangGraph 1.2.11LangGraph 1.2.11 (checkpoint 4.2.0, checkpoint-sqlite 3.1.1) could not be installed — outbound PyPI DNS was unavailable — so the round pins public release/source contracts and compares them with AER-0's executable invariants. Runtime-owned persistent state, checkpoint history, version tracking, long-term stores and TTL memory all converge (LangGraph is richer on resume, time travel, concurrency and delta checkpointing). What remains AER-default-distinct: candidate → verify → commit, mandatory fact provenance, node-level optimistic epistemic commit, task-conditioned tension refresh, capability/container separation, explicit epistemic-operator routing, and the graph meaning an epistemic world representation rather than control flow.
- ExperimentEXP-2026-0004R3 — epistemic commit transaction vs scattered and centralized application gatesAll canonical writes in AER-0 were migrated to an Epistemic Commit Transaction (claim, evidence, provenance, valid/observed time, expected version, operator, verification policy, authority, dependencies, refresh policy); direct candidate commit became a negative path. Three systems ran the same governance witnesses: an application whose policy is scattered across callers, an application with one centralized mandatory gate, and AER-ECT. The centralized application gate matched AER-ECT on every tested property; computational uniqueness is not supported. 48 tests.
- ExperimentEXP-2026-0005R4 — policy mutation surface: scattered governance vs one mandatory boundaryFor N ∈ {1, 4, 16, 64} canonical-write callers and R = 6 governance rules, policy sites and rule placements grow as N and N·R under scattered governance versus 1 and R under either a central application gate or AER-ECT; a new rule needs N caller edits versus one; a 75 % caller-local migration leaves 16 of 64 callers on the old contract. The counterweight is explicit: one omitted local rule has blast radius 1/N, a defect in the shared gate has blast radius 1, and under an equal-p toy omission model the expected exposed-caller fraction is identical for all three systems.
- ExperimentEXP-2026-0006R5 — complete mediation and bypass resistanceAn adversarial boundary test across five threat layers, after adding a SQLite authorizer on managed connections, trigger guards on the five protected tables, verifier-only verification-state recording, BEGIN IMMEDIATE commit serialization, and reference-monitor / canonical-integrity audits. Ten scenarios: five PREVENTED (managed direct write, foreign raw write with intact schema, post-verification tamper, fabricated verdict, managed guard drop), three OPEN_DETECTED (foreign guard drop, post-guard-removal node forge, hostile same-process disable), one OPEN_UNDETECTED (full-DB writer forging node and backing candidate consistently), one same-base two-process race SERIALIZED_TO_SEMANTIC_CONFLICT.
- ExperimentEXP-2026-0007R6 — external trust anchor, process-separated writer, adapter conformanceEvery externally anchored commit binds a digest over the accepted transition into an Ed25519-signed hash chain with an optional out-of-band latest-head receipt; signing authority moves into a child writer process that never returns the private key; a behavioural conformance harness defines what any future state adapter must pass. The R5 coherent full-DB forgery stays OPEN_UNDETECTED for the internal audit and becomes OPEN_DETECTED under the external anchor audit; anchor mutation is DETECTED; valid-prefix truncation is undetected without a trusted head and detected with one; a required-anchor failure fails closed on the tested path; an orphan anchor is detected; the SQLite adapter passes the contract and a deliberately broken adapter is rejected. 79 tests.
- ExperimentEXP-2026-0008PACC-Lab v0.1 — E0–E4: first micro-witnessThree primitive-level non-probabilistic systems (N0 signed support, N1 ordinal tournament, N2 signed graph) integrate evidence over four hidden hypotheses next to an exact Bayesian reference. All three pass the Level-3 micro criteria — held-out JS ≤ 0.0101, update JS ≤ 0.0042, intervention JS ≤ 0.0058, action agreement ≥ 0.957 — in both uniform and heterogeneous evidence quality, with shuffled-target controls around 0.31–0.33. The redundancy diagnostic collapses N0 and N1 into one exact reparameterization family, leaving two independent families against a PACC-A minimum of three.
- ExperimentEXP-2026-0103XA-06 first real-model pilot — Qwythos-9B-v2 on the A0→A5 ladder (36 trials, 2026-09-03)The first time a real model was placed inside the IPM instrument: hf.co/empero-ai/Qwythos-9B-v2-GGUF:Q4_K_M served by Ollama on an RTX 3070, run locally on 2026-09-03 by Splice (Claude Code) on Neo.K's authorization through XA-06L — MATH-003, CODE-001, CON-003 × A0–A5 × 2 replicates, 36 trials, 186 invocations, 162 trajectories, telemetry complete on every trial. Sealed REAL_MODEL_PILOT_INCOMPLETE: 33/36 complete, three trials aborted because the same-model verifier returned non-JSON to a strict parser (a real model behaviour, deliberately not re-rolled). SSR = 1.0, SDR = 0.0: scaffolding brought no measured quality gain on these three easy tasks while A5 used 3.10× the device energy and 3.13× the wall time of A0, and the fixed eight-sample conditions A2–A4 used 7.9–9.3×. The apparent A3/A4 quality rise to 0.667 is a missingness artifact; and for two of the three tasks the recorded 'quality' was output-format compliance, not task correctness (post-hoc: 33/33 completed outputs semantically correct; strict output-contract compliance 12/36).
- ExperimentEXP-2026-0102XA-05 — 36-trial end-to-end smoke gate with a scripted provider (synthetic)Runs MATH-003, CODE-001 and CON-003 × A0–A5 × 2 replicates through XA-04 + XA-03 + XA-02 with a ScriptedProvider whose outputs were constructed so the expected quality curve is known in advance (A0 0.30, A1 0.633, A2–A5 1.0 → SSR 0.30 / SDR 0.70 as fixtures). 36 trials, 180 invocations, 168 trajectories, 6 retries, 6 tool calls, 24 verifier passes; candidates 168 = 36 selected + 132 discarded; 0 accounting mismatches, 0 private-reference leaks; energy null on purpose (NullCollector) to prove Unknown ≠ 0. The package states scientific_interpretation_allowed: false.
- ExperimentEXP-2026-0101Experiment A — single-pass vs scaffolded controlled measurement (protocol v0.1)The first v0.2 experiment: for one model and a fixed task set, climb the scaffold ladder A0 native single pass → A1 extended trajectory → A2 multi-sample → A3 verifier → A4 deterministic local tools → A5 bounded full agentic loop, recording quality and physical cost at every level to obtain a scaffolding response curve, SSR/SDR, scaffold cost multipliers and marginal yields. Five hypotheses (H1 gain exists, H2 gain has physical cost, H3 marginal yield is non-constant, H4 native and system capability are distinguishable, H5 models have different scaffolding profiles), pre-registered quality projections, initial-information equality across A0–A3, budget caps, n = 5 pilot / 20 formal replicates, failure classification and the rule that a null result is not an experiment failure. Status READY FOR PILOT; the thirty-task run itself has not been executed — only the three-task instrument gates below.
- ExperimentEXP-2026-0104Experiment B — binary vs numeric human measurement (declared)Compare direct 0–10 rating, structured yes/no items and adaptive pairwise comparison on response time, missingness, inconsistency, test–retest, predictive validity and fatigue, with participants randomized across formats. Declared in the canonical index as the third v0.2 experiment; not designed in detail and not run.
- ExperimentEXP-2026-0105Experiment C — μI operational identification (declared)On proof steps, code repair and constraint puzzles, construct candidate semantic transitions, then test them by ablation and counterfactual replacement to see whether an effective semantic count N_μ can be identified and whether it predicts anything. Declared as the fourth v0.2 experiment; not run.
- ExperimentEXP-2026-0106Experiment D — physical trace alignment (declared)On one machine, align GPU power telemetry, latency, peak memory, memory bandwidth and device occupancy with execution events — the telemetry side that XA-03 already provides — before any claim about data-center energy. Declared as the second v0.2 experiment in the recommended order; the pilot's XA-03 traces are its first raw material, but the alignment study itself has not been run.
- ExperimentEXP-2026-0107Experiment E — token / FLOPs proxy failure test (declared)Executions of the same task quality under different languages, verbosity, context lengths and memory pressure, comparing token count, FLOPs, energy, time, memory traffic and N_μ^eff to see where token and FLOPs stop tracking cost and work. Declared as the last v0.2 experiment; not run.