EveMissLab
AI Research Laboratory
Theory, experiments, models, data and evidence for AI systems.
EveMissLab studies AI as a computational, architectural, representational and autonomous research subject — from theory to reproducible evidence.
- Active research
- 7
- Active experiments
- 1
- Datasets
- 2
- Benchmarks
- 4
- Programs
- 2
Snapshot AI-SNAPSHOT-v0.1-fe85b9694a45 · 124 research objects · 499 relations · generated 2026-09-11 · index.json
Current research
View all (7) →- ResearchRES-2026-0002PACC conjecture — do non-probabilistic primitives converge to probability-like structure?Systems whose canonical state and update rules never require probability distributions, Bayesian posteriors or sampling are placed in the same evidence-integration tasks as an exact Bayesian reference. The lab measures whether low-complexity train-only maps carry their states to the Bayesian state, whether the maps commute with updates and survive interventions, and whether independently designed families converge — a four-level ladder (PACC-B, -R, -D, -A) with preregistered falsification conditions.
- ResearchRES-2026-0004A PACC-style runtime on language models: reasoning, intent and creative breadthDoes wrapping candidate selection in a canonical hard/derived-constraint commit space change what a generator produces? A synthetic A/B/C witness (generator only / hard verifier / PACC runtime) shows derived coherence and constraint-satisfying novelty up, soft-preference fit slightly down, and a creative-breadth collapse that a post-hoc 'elastic' exploration policy recovers. The real-language-model version of the benchmark has a validated, fail-closed harness but has not been executed.
- ResearchRES-2026-0001Blind derivation of an adaptive epistemic architectureStarting only from first principles — temporally heterogeneous world knowledge, canonical symbolic state, adaptive representation space, capability memory, substrate-neutral compute — the eleven-paper series derives a candidate architecture and then asks the uncomfortable question the derivation raised: why does it look so much like modern compound AI, and does any of the difference survive into a runtime?
- ResearchRES-2026-0003Does epistemic governance survive implementation? The AER-0 architecture comparisonSix rounds compare the AER-0 runtime with progressively stronger baselines — a stateless recomputer, fixed-TTL state, memory + tools, an evented compound agent, LangGraph 1.2.11 (source-grounded), and finally a centralized application-level gate that is allowed to be as good as AER. Each round narrows the claim: from 'AER is distinct' to 'AER promotes epistemic governance from application convention to a runtime-level canonical-write boundary', with bounded mediation and an external signed authenticity witness, and explicitly no claim of unique computational capability.
- ResearchRES-2026-0103Capability line — how much intelligence remains without the loop, and the unified intelligence eventPapers 09–10 and the v0.2 Experiment A instrument. A system is (model, scaffolding vector); single-pass quality Q_SP and full-system quality Q_F give the scaffolding survival ratio SSR = Q_SP/Q_F, the dependence ratio SDR, the scaffold cost multiplier SCM and marginal scaffolding yields along a controlled ablation ladder A0…A5 — all read together with the physical overhead, because a loop is not cheating; hiding its cost is. Paper 10 packs quality, semantic work, physical computation and scaffolding into the canonical intelligence event, compares systems on Pareto frontiers under a no-premature-scalarization rule, and fixes a minimum reporting standard. The line's first real data point is the 2026-09-03 pilot on a local 9B model.
- ResearchRES-2026-0101Execution and physical line — what one answer costs in turns, semantic work, energy and computational spacetimePapers 01–05. Decomposes the chat-interface 'turn' into user turns, generation trajectories, model invocations, external loops, retries, selection and a physical execution trace; proposes the minimum intelligent semantic execution unit μI as the missing middle layer between physical primitives and task achievement; borrows the cross-level, latent-inference, population-coding and causal-perturbation method of neuroscience; types energy as gross / baseline / marginal / attributed with a declared boundary and keeps Landauer as a bound, not a price; and replaces FLOPs with a physical cost vector and a vector-first computational spacetime with its own topology.
Theory
View all (17) →- TheoryTHY-2026-0005PACC conjecture — the four-level convergence ladderPACC-B (behavioural: D_B ≤ ε), PACC-R (representational: a low-complexity map Φ: S_N → S_P valid on held-out tasks, K(Φ) ≤ κ), PACC-D (dynamical: Φ∘U_N ≈ U_P∘Φ) and PACC-A (architecture attractor across ≥3 independent designs). The conjecture pre-registers its own failure modes F1–F6 — behavioural divergence, no low-complexity map, update diagram fails, intervention divergence, architecture diverges under scale, probability-specific advantage persists — so that 'similar' and 'probability' cannot be redefined after the fact.
- TheoryTHY-2026-0001Asymmetric spacetime tension: temporally heterogeneous world knowledgeAsymmetry is lifted from edge direction or weight to the effective time scale of nodes and relations. Each node carries stability, information-decay rate, update tension, system impact and local valid time; refresh is triggered by tension, external disturbance, information age, change velocity and event relevance rather than by one global clock, so stable knowledge sleeps, volatile knowledge refreshes often, and a rarely changing high-impact node triggers wide dependency recomputation when it does change.
- TheoryTHY-2026-0002Canonical symbolic state and candidate → verify → commit authorityNatural language is input and output, never the canonical state. Text is parsed, normalized, semantically bound and source-tagged into comparable, verifiable symbolic structure; output is rendered from that state. Model and tool outputs are candidate evidence, and only a committer, after a passing verifier and an expected-version check, may mutate canonical state.
- TheoryTHY-2026-0003Intelligent architecture attractorDifferent first principles may compile into the same small family of computational forms. The comparison framework separates six levels of difference — code, primitive, computation trace, architectural state semantics, observable behaviour, resource efficiency — and asks who owns state, what may rewrite canonical state, and whether dynamics are preserved under low-cost mapping. The series ends with a fixed verdict map: Distinct Advantage, Operational Convergence, Behavioral Equivalence Only, Inconclusive, Architecture Worse.
- TheoryTHY-2026-0004Probability is not Bayesian; Bayes cannot self-authorize its premisesA system can be stochastic, probabilistic or probability-shaped without performing Bayesian conditionalization, and can look Bayesian without a real prior, likelihood or posterior. A layered vocabulary (stochastic, probabilistic, Bayesian-like, exact, approximate, generalized Bayesian) and a Bayesian authenticity test check whether an update is substantively Bayesian or merely redescribed as such; and the update rule itself — prior, likelihood, hypothesis space — needs a justification that Bayes' rule does not supply (Papers 09–10).
- TheoryTHY-2026-0006AER-ECT: a mandatory epistemic commit transaction boundaryEvery canonical knowledge mutation must pass one mandatory epistemic transaction boundary binding claim, evidence, provenance, valid and observed time, expected version, epistemic operator, verification policy, authority, dependencies and refresh policy: Propose → Verify → Authorize → CompareAndCommit → BindDependencies → Audit. Its parts are not new (truth maintenance, transaction logic, PROV, bitemporal state, PDP/PEP); the hypothesis is that the tuple is mandatory at the canonical write boundary — an epistemic reference-monitor-like boundary, not a proven tamperproof monitor.
Claims
View all (5) →- ClaimCLM-2026-0101F1 — Token hypothesisIf the token is a good universal unit of intelligent work, then N_μ^eff / TokenCount should be approximately stable across models, languages and phrasings once task quality is controlled. Large drift across models, modalities or expression forms weakens 'token = intelligent work unit' to an implementation/interface proxy.
- ClaimCLM-2026-0102F2 — FLOPs sufficiencyIf FLOPs suffice to describe physical computational cost, then with operations controlled, wall time, energy, memory traffic, interconnect traffic and memory residency should not show large independent variation. If same-FLOPs workloads differ greatly because of memory patterns or topology, a single FLOPs cost model is insufficient.
- ClaimCLM-2026-0103F3 — Binary burden hypothesisIf direct numeric human rating were already the minimum-burden, high-quality measurement interface, then well-designed binary / pairwise adaptive protocols should have no advantage in response time, consistency, dropout, predictive validity or fatigue. BRQM predicts an advantage in at least some settings; it can be tested by randomizing participants across direct 0–10, structured yes/no and adaptive pairwise formats.
- ClaimCLM-2026-0104F4 — Scaffolding separationIf scaffolding does not change the structure of the capability source, SSR ≈ 1 should hold across most tasks and test-time compute budgets. If a large share of benchmark quality appears only under multi-sample, tool, verifier or loop conditions, the distinction between model-native and system capability is empirically necessary. First data point: three easy tasks on a local 9B model gave SSR = 1.0 — consistent with 'no gap' on tasks the native pass already solves, and uninformative beyond that because the quality axis was confounded by format compliance.
- ClaimCLM-2026-0105F5 — Semantic intermediate utilityIf the semantic middle layer N_μ has no measurement value, then adding it should not improve the explanation or prediction of cross-architecture efficiency, error paths, scaffold gain or task transfer over a direct physical-cost → quality model. μI itself must pass this predictive/explanatory utility test or be revised or eliminated.
Active experiments
View all (1) →Latest results
View all (18) →- ResultRST-2026-0014Real local-model A/B/C table (Qwythos-9B-v2)The v0.2 protocol executed for the first time on a real language model — a local open-weight 9B (Qwythos-9B-v2, Q4_K_M, Ollama, thinking off) as generator, selector and judge; 16 tasks × 2 repetitions × 4 candidates, 380 calls, 32 rows per condition. The PACC runtime (C) gains on supersession alignment (+0.031 vs B, +0.097 vs A) and repair success (+0.070 / +0.094), the two axes the canonical-intent compilation step exists for; derived coherence (+0.011) and intent persistence (−0.002) do not move; valid novelty is slightly lower (−0.045 vs B). Most per-task pairs are ties because the 9B judge saturates near 1.0, and with two repetitions per task the judge's free-text pattern labels never repeat, so creative breadth is not measurable. The conditions chose different candidates in 69 % of task × repetition pairs.
- ResultRST-2026-0015Run 2: no reliable condition difference; breadth not reduced (64 rows)C vs B: supersession +0.0005, repair -0.0125, derived coherence -0.0075, intent -0.0114, valid novelty -0.0187; label-free breadth: cluster entropy C 0.923 / A 0.909 / B 0.864.
- ResultRST-2026-0016Run 3: valid novelty and repair up for the PACC runtime with thinking on; governance flat; breadth not reduced (32 rows)C vs B: valid novelty +0.0687, repair +0.0153, semantic novelty +0.0166, derived coherence -0.0063, intent -0.0053, supersession -0.0059; C vs A: valid novelty +0.0375, repair +0.0466; label-free breadth ratio C 1.047 / A 0.960 / B 0.973.
- ResultRST-2026-0007v0.2 adversarial geometry: N3 converges, N2 crosses the representation gateN3 held-out JS 0.014454 (Level 3); N2 held-out JS 0.053327 > 0.05 (Level 1); N0/N1 0.023921 (Level 3).
- ResultRST-2026-0008v0.4: latent dependence coordinate not recoveredDependence-coordinate pass count 0 for all four systems in every seed; in high correlation real JS ≈ 0.0398 is below the 0.05 absolute gate but no better than shuffled (≈ 0.0400) or constant (≈ 0.0397).
- ResultRST-2026-0009v0.6: composed coordinate 4/4, direct 0/4Composed latent JS ≪ shuffled/constant and ≪ broken composition for every family and geometry (e.g. 0.003603 vs 0.061508 vs 0.058196, moderate N0).
Data & benchmarks
Data · 2
- DatasetDAT-2026-0001PACC synthetic evidence worldsDeterministically generated micro-environments: 4 hidden hypotheses × 6 binary prototypes with per-feature reliabilities; source-quality worlds for learned reliability (20 worlds / 280 episodes / 28 observations at primary scale); correlated-cluster worlds with latent shared inversion at q ∈ {0.06 … 0.42}; and regime-switching dynamic worlds. Primary seeds 20260908 (v0.1) and 20260909 (v0.3+); secondary seeds are labelled post-hoc.
- DatasetDAT-2026-0002PACC-Hybrid v0.1 shared candidate pools1,568 candidate-pool instances (7 categories × 28 tasks × 8 repetitions, 112 candidates each) generated synthetically and given identically to all A/B/C conditions, with primary, multiseed and elastic-diagnostic result JSON. No language-model output is included.
Benchmarks · 4
- BenchmarkBEN-2026-0001PACC micro-lab protocol v0.1 (frozen gates)The preregistered decision rule reused unchanged from v0.1 to v0.13: E0 purity, held-out representation distance D_R = E[JS(Φ(S_N), S_P)] ≤ 0.05, update commutation D_U ≤ 0.05, off-manifold intervention D_I ≤ 0.08, action agreement ≥ 0.85, a shuffled-target negative control the real map must beat, a design-independence rule (agreement ≥ 0.999 and R² ≥ 0.995 collapse two designs into one family), verdict levels 0–3, and a strong PACC-A gate of at least three independent convergent families.
- BenchmarkBEN-2026-0003PACC-LLM Hybrid A/B/C benchmarkThree conditions over identical candidate pools: A generator-only, B hard verifier, C PACC runtime (plus post-hoc D 'elastic'). v0.1: 7 categories × 28 tasks × 8 pools × 112 candidates (1,568 shared pool instances). Metrics: explicit hard adherence, derived coherence, soft-intent satisfaction, long-horizon retention, raw and valid novelty, pattern entropy. v0.2: 16 hand-authored natural-language tasks, equal accounted budgets, gold rubric visible only to a condition-blind judge.
- BenchmarkBEN-2026-0002AER-0 architecture-comparison suite (R1–R6)Deterministic, executable comparisons that grow round by round: R1 — two facts (one stable, one changing at hours 8 and 16), eight queries over 17 simulated hours, baselines B1 stateless / B2 fixed 6-hour TTL / B3 memory + tools / B4 evented compound agent; R3 — scattered policy vs centralized application gate vs AER-ECT on provenance, stale-write, audit and dependency binding; R4 — policy-topology scaling over N ∈ {1, 4, 16, 64} callers × 6 rules; R5 — ten bypass scenarios across five threat layers; R6 — coherent full-DB forgery, anchor mutation, prefix truncation with/without a trusted head, fail-closed anchor, orphan anchor, adapter conformance.
- BenchmarkBEN-2026-0101XA-02 — 30-task pilot pack for the A0→A5 scaffolding response30 tasks — 10 math, 10 code, 10 constraint — with a public task file (prompt, output contract, pre-registered quality projection), a private reference file that must never enter model context, a deterministic evaluator and a 30/30 self-test. Math and constraint answers are one JSON object; code answers are Python source scored by hidden tests, and code execution is refused unless explicitly enabled inside an external sandbox. The pack's own words: not a general intelligence benchmark but a controlled instrument for measuring scaffolding response under Experiment A.
Research programs
View all (2) →- ProgramPRG-2026-0001Adaptive Epistemic SystemsA research program that derives an AI runtime from first principles — world knowledge changes at heterogeneous rates, natural language is a rendering rather than the canonical state, capability must be remembered and reused, algorithm and compute substrate are separate — and then forces the derivation into executable reality: the AER-0 reference runtime with six architecture-comparison rounds, and the PACC micro-lab on whether non-probabilistic primitives converge to probability-like structure.
- ProgramPRG-2026-0101Intelligence Physical Metrology (IPM)A research program that asks not how smart an AI is but what one answer costs a physical system: how many user turns, generation trajectories and hidden loops it took, how much effective semantic work was done, how much energy and computational spacetime was occupied, how much external scaffolding was leaned on, and how much verifiable quality came back. Ten theoretical papers (2026-09-02) define the measurement objects — task, quality, semantic work, physical computation, scaffolding capability, measurement metadata — and five falsifiable propositions; the v0.2 experimental protocol (Experiment A, single-pass vs scaffolded) is instantiated as the XA-02…XA-06L instrument packages and was executed once on a local 9B model.
Models & architectures
View all (6) →- ModelMOD-2026-0006Qwythos-9B-v2 (Q4_K_M, local, Ollama)Open-weight 9.0B model of the qwen35 family, GGUF Q4_K_M, served locally by Ollama 0.33.3 on an RTX 3070. Used as generator, selector and judge in the real local run of the PACC-Hybrid v0.2 protocol, with thinking disabled in runs 1–2 and enabled (effort medium, +3,500 output tokens per call) in run 3.
- ModelMOD-2026-0001N0 — signed supportAdditive signed evidence support per hypothesis; under uniform reliability it is closely related to rescaled log-evidence accumulation, which is why the heterogeneous and adversarial stresses exist.
- ModelMOD-2026-0002N1 — ordinal tournamentPairwise ordinal competition between hypotheses; behaviourally and linearly identical to N0 on every tested trajectory.
- ModelMOD-2026-0003N2 — signed graphSigned relation graph over hypotheses with graph dynamics; the basin-boundary family — fails the representation threshold under adversarial geometry (v0.2), is less stable under high correlation (v0.4), and has the narrowest, seed-sensitive transfer basin (v0.8).
- ModelMOD-2026-0004N3 — constraint competitionA nonlinear constraint-field system added in v0.2 as the third structurally distinct family; converges bidirectionally across geometries (v0.7) and shares the wider N0/N1 basin (v0.8).
- ModelMOD-2026-0005Bayesian reference (exact posterior / Beta-Bernoulli / joint common-cause / HMM)The probabilistic side of every comparison: exact posteriors with numerical reliabilities, Beta-Bernoulli source-quality learning, the joint-likelihood common-cause model for correlated clusters (and its naive-independence negative control), and a Bayesian HMM for regime switches.
Latest papers
View all (24) →- PaperPAP-2026-0009Paper 09 — Probability is not BayesianA layered vocabulary from stochastic to generalized Bayesian and a Bayesian authenticity test for updates that are only redescribed as Bayesian.
- PaperPAP-2026-0005Paper 05 — Memory, algorithm libraries and reusable solution pathsAlgorithms, tools, contracts, costs and outcome history become capability nodes; Reuse ≻ Adapt ≻ Create.
- PaperPAP-2026-0006Paper 06 — Substrate-neutral computational containersAny execution environment with representable input, valid transition and readable output is a compute container; algorithms and containers are scheduled jointly.
- PaperPAP-2026-0007Paper 07 — Blind-derived AI: do different first principles converge on one engineering form?Treats the first six papers as a blind-derivation experiment and proposes a protocol and five levels of similarity for comparing the result with modern AI.
- PaperPAP-2026-0008Paper 08 — Intelligent architecture attractors: at which level does difference live?Six levels of architectural difference, architecture equivalence classes and attractor basins; the executed architecture is the architecture.
- PaperPAP-2026-0010Paper 10 — Bayes within Bayes: who authorizes the update rule?Prior, likelihood and hypothesis space need an authorization Bayes' rule cannot give; meta-epistemic configuration becomes part of the architecture.
Research systems
View all (7) →- SystemSYS-2026-0002PACC-Lab — micro-lab harness for the convergence conjectureFully observable evidence-integration environments, four primitive-level non-probabilistic families (N0–N3), exact Bayesian references, train-only low-complexity mappers, shuffled/constant/broken-composition controls, a design-redundancy diagnostic and frozen preregistered thresholds carried unchanged from v0.1 through v0.13. Each release is a sealed FINAL bundle with its protocol, results and next-experiment preregistration.
- SystemSYS-2026-0003PACC-LLM Hybrid LabAn A/B/C harness — generator only, hard verifier, PACC runtime — over shared candidate pools with equal accounted budgets. v0.1 is a synthetic architecture witness (zero LLM calls); v0.2 adds hand-authored natural-language tasks, a condition-blind judge, literal machine checks independent of the judge, fail-closed behaviour without an API key, and frozen-cache replay for exact reproduction after a live run.
- SystemSYS-2026-0001AER-0 — Adaptive Epistemic Runtime MVPA Python + SQLite reference runtime that owns persistent canonical state; a language model is optional and replaceable. Candidate → verify → versioned commit, mandatory provenance for facts and relations, freshness and confidence as separate state, an asymmetric tension scheduler with shock-triggered refresh, capability/container registries, persisted workflow memory (reuse/adapt/create), and a deterministic/rule/Bayesian/heuristic epistemic router. Grown through six comparison rounds into an epistemic commit transaction with SQLite-level mediation, an Ed25519-signed external anchor and a process-separated writer (79 tests at R6). Developed under mssp-tdd-apr as a DEGRADED-TWIN run — no independent twin verification is claimed.
- SystemSYS-2026-0101XA-03 — telemetry and run logger (physical execution evidence layer)Provider-agnostic logger that records one benchmark run as append-only events plus telemetry and derives a typed summary deterministically: model-invocation, trajectory, tool-call and verifier spans, candidate created/discarded accounting, wall time, device-time, GPU power integrated to device_energy_j (energy type device_measured, never relabelled as marginal), peak memory and memory residency, with unknowns kept null (Unknown ≠ 0) and forbidden conversions (tokens → J, TDP → J, price → J, GPU-hours → J). Golden fixture 450 J / 1.75 util·s / 12 GB peak / 33 GB·s residency verified; 35 passed, 1 skipped in 0.39s; instrumentation burden is calibrated and reported, not subtracted.
- SystemSYS-2026-0102XA-04 — A0→A5 model runner and scaffold orchestratorA policy-driven state machine that runs the same TrialExecutor under immutable condition policies: A0 native single pass, A1 extended single trajectory (2× generation budget), A2 eight samples with exact-majority selection, A3 eight samples plus a typed verifier, A4 verifier plus a bounded deterministic tool loop, A5 a bounded PLAN → ACT → OBSERVE → VERIFY → REVISE loop. XA-02 and XA-03 are external canonical dependencies (declared by hash, not vendored); tool actions under A0–A3 are protocol violations; benchmark scoring happens only after the XA-03 run is finalized so scoring cost is never charged to the system; private references are never read by the task loader; no private chain-of-thought is captured. 29 tests, status CANDIDATE_CLEAN_VERIFIED.
- SystemSYS-2026-0103XA-06 — real-model pilot gatePuts a real model inside the validated XA-01→XA-04 instrument: preflight (one isolated MATH-003/A0 trial through provider, XA-03 and XA-02), the canonical 36-trial matrix (MATH-003, CODE-001, CON-003 × A0…A5 × 2 replicates), gate verification and analysis. Frozen temperature / top-p / seed, the A0/A1 generation-budget multiplier forwarded into the token budget, a calculator tool contract injected only for A4/A5 generator requests, disabled code evaluation treated as quality-unavailable rather than measured zero, and the rule that engineering validation can never be promoted to pilot completion: 'no real model run, no real model claim'. Valid real outcomes explicitly include A0 = A5, A3 < A2, A5 < A0 and SSR = 1. 20 tests; the package itself contains no pilot result (status READY_FOR_REAL_MODEL).
Research graph
View all (499) →Archive
View all (0) →No public records yet.
Status and evidence vocabulary
- Research status
- IDEA
- only a question or a concept so far
- PRELIMINARY
- a first formalisation or observation exists
- ACTIVE
- under continuous research
- EXPERIMENTAL
- being tested experimentally
- VALIDATING
- the main proposition is formed; evidence is being confirmed
- REPLICATING
- being reproduced, or checked across models
- STABLE
- the current conclusions are relatively stable
- PAUSED
- paused
- ARCHIVED
- no longer actively pursued
- SUPERSEDED
- replaced by a newer version or theory
- Evidence level
- E0
- Concept only
- E1
- Internal observation
- E2
- Controlled experiment
- E3
- Repeated experiment
- E4
- Cross-model / cross-environment replication
- E5
- External reproduction / independent evidence
- Data basis
- THEORY
- Theoretical reasoning only; no measurement.
- SYNTHETIC
- Synthetic data and theoretical reasoning. Many now treat synthetic data as if it were real; this laboratory says the opposite deliberately — until a real hybrid model exists, an inference is only an inference, and theoretically possible is not actually possible.
- DETERMINISTIC RUNTIME
- A deterministic scenario executed in a real runtime; the code and procedure are listed, and the numbers are what they are.
- REAL MODEL
- A real language model was executed; the model, its version and its configuration are recorded on the page.
- NOT RUN
- Designed and validated as a harness; never executed with a real model.
Research status says where a line is in its life; evidence level says how much has been shown. They are recorded separately on purpose: ACTIVE or STABLE is not a truth claim, and an ARCHIVED line can still carry E4 evidence.
Data basis says what a number was measured on. Synthetic results are labelled SYNTHETIC on every card and page: they are evidence about the architecture under synthetic conditions, not about the world.