AI Research Laboratory17 records
Theory
Formal claims, each with its assumptions, predictions and the conditions under which it fails.
- TheoryTHY-2026-0005PACC conjecture — the four-level convergence ladderPACC-B (behavioural: D_B ≤ ε), PACC-R (representational: a low-complexity map Φ: S_N → S_P valid on held-out tasks, K(Φ) ≤ κ), PACC-D (dynamical: Φ∘U_N ≈ U_P∘Φ) and PACC-A (architecture attractor across ≥3 independent designs). The conjecture pre-registers its own failure modes F1–F6 — behavioural divergence, no low-complexity map, update diagram fails, intervention divergence, architecture diverges under scale, probability-specific advantage persists — so that 'similar' and 'probability' cannot be redefined after the fact.
- TheoryTHY-2026-0001Asymmetric spacetime tension: temporally heterogeneous world knowledgeAsymmetry is lifted from edge direction or weight to the effective time scale of nodes and relations. Each node carries stability, information-decay rate, update tension, system impact and local valid time; refresh is triggered by tension, external disturbance, information age, change velocity and event relevance rather than by one global clock, so stable knowledge sleeps, volatile knowledge refreshes often, and a rarely changing high-impact node triggers wide dependency recomputation when it does change.
- TheoryTHY-2026-0002Canonical symbolic state and candidate → verify → commit authorityNatural language is input and output, never the canonical state. Text is parsed, normalized, semantically bound and source-tagged into comparable, verifiable symbolic structure; output is rendered from that state. Model and tool outputs are candidate evidence, and only a committer, after a passing verifier and an expected-version check, may mutate canonical state.
- TheoryTHY-2026-0003Intelligent architecture attractorDifferent first principles may compile into the same small family of computational forms. The comparison framework separates six levels of difference — code, primitive, computation trace, architectural state semantics, observable behaviour, resource efficiency — and asks who owns state, what may rewrite canonical state, and whether dynamics are preserved under low-cost mapping. The series ends with a fixed verdict map: Distinct Advantage, Operational Convergence, Behavioral Equivalence Only, Inconclusive, Architecture Worse.
- TheoryTHY-2026-0004Probability is not Bayesian; Bayes cannot self-authorize its premisesA system can be stochastic, probabilistic or probability-shaped without performing Bayesian conditionalization, and can look Bayesian without a real prior, likelihood or posterior. A layered vocabulary (stochastic, probabilistic, Bayesian-like, exact, approximate, generalized Bayesian) and a Bayesian authenticity test check whether an update is substantively Bayesian or merely redescribed as such; and the update rule itself — prior, likelihood, hypothesis space — needs a justification that Bayes' rule does not supply (Papers 09–10).
- TheoryTHY-2026-0006AER-ECT: a mandatory epistemic commit transaction boundaryEvery canonical knowledge mutation must pass one mandatory epistemic transaction boundary binding claim, evidence, provenance, valid and observed time, expected version, epistemic operator, verification policy, authority, dependencies and refresh policy: Propose → Verify → Authorize → CompareAndCommit → BindDependencies → Audit. Its parts are not new (truth maintenance, transaction logic, PROV, bitemporal state, PDP/PEP); the hypothesis is that the tuple is mandatory at the canonical write boundary — an epistemic reference-monitor-like boundary, not a proven tamperproof monitor.
- TheoryTHY-2026-0007Capability memory and substrate-neutral compute: Reuse ≻ Adapt ≻ CreateBeyond facts, the system keeps the algorithms, tools, execution contracts, costs, versions, applicability conditions and success/failure history it has used, and prefers reusing a known solution path over adapting one over creating one. Algorithms and compute containers are different layers: any environment that accepts representable input, performs a valid state transition and returns readable output is a container with its own cost, latency, error and availability model, so algorithm/container pairs are selected jointly.
- TheoryTHY-2026-0101Turn decomposition and externally loopless intelligenceA 'turn' is decomposed into the vector T = (U user turns, G generation trajectories, I model invocations, L external feedback loops, R retries, S selection/verification, P physical execution trace). The cleanest single pass is U=1, G=1, R=1, L=0, S=0, and it excludes only action→new-evidence→replanning, not autoregressive sequential computation. Loops come in five kinds (tool, environment, verifier, retry, candidate/selection), capability lives in four layers (intrinsic, elicited, system, product), and every 'one-turn' claim projects onto the event vector (Q, U, G, I, L, R, S, T, E, V_CST).
- TheoryTHY-2026-0102μI — the minimum intelligent semantic execution unit (candidate theory)μI is a task-relative semantic state transition z_t → z_{t+1} that satisfies seven conditions: task relevance, state change, causal contribution (Q(Y | μ) > Q(Y | do(μ=0))), non-decomposability at the chosen resolution, composability, realization independence and physical realizability. It is neither token, FLOP, neuron activation, layer nor thought; its minimality is resolution-, task- and observer-relative. Four candidate families (belief update, relation construction, constraint resolution, information gain) unify as a task-relevant, causally useful, non-redundant Δz. Gross vs effective counts give the semantic efficiency η_μ, and a cross-level map μI → ρ_C → ρ_P → ρ_T leads down to physical cost.
- TheoryTHY-2026-0103Cross-level triangulation and measurement gradesNeuroscience has no '1 thought = N spikes' conversion; what it has is a cross-level proxy method: behavior → latent cognitive model → neural code → cellular events → physical implementation, each layer with its own units. IPM borrows five principles (level separation, latent inference, population over atomism, encoding–decoding duality, causal perturbation), extends Marr's three levels to five (task, semantic, algorithmic, physical events, thermodynamic), and defines cross-level triangulation for AI as E = (output, semantic, internal trace, ablation, hardware) evidence with a graded confidence in μI from D (behavioral) to A+ (physical-semantic alignment).
- TheoryTHY-2026-0104Energy accounting hierarchy and thermodynamic type safetyNeural energetics builds energy bottom-up (membrane dynamics → ion flux → pump work → ATP → dissipation) and finds that a spike has no fixed energy and that most cortical signaling energy is spent on synaptic integration and state maintenance, not the visible pulse. IPM copies the discipline, not the numbers: energy is typed as E = (gross, baseline, marginal, attributed, thermodynamic minimum), a boundary and baseline rule must be declared, information per Joule is not intelligence per Joule, Landauer's kT ln 2 bounds erasure and is not the price of a μI, and Shannon or variational 'energies' never become physical Joules without an explicit mapping.
- TheoryTHY-2026-0105Physical computation cost vector and computational spacetimePhysical cost is the vector C_P = (typed operations, memory traffic by hierarchy, I/O, interconnect, memory residency, device occupancy, wall time, energy), not FLOPs. Computational spacetime is first a measure V_CST = ∫ R(t) dt = (V_C, V_M, V_N, V_S) over a resource field; the components may not be added before a declared normalization, and equal volume (8 GPU × 10 s = 1 GPU × 80 s) is not equal topology Θ_CST = (T_wall, T_serial, P_parallel, D_peak, M_peak, B_peak, Γ_comm). Roofline, memory-wall and data-movement results explain why same-FLOPs workloads differ in time and energy; a peak hardware footprint is a capacity barrier; CST grades run from D (spec estimate) to A+ (causal resource attribution).
- TheoryTHY-2026-0106Structured quality, hard gates and the specification–verification separationQuality is Q(Y | X, S, W, B_Q): output relative to task, specification, environment and boundary. It is measured as a structured vector (correctness, alignment, completeness, consistency, robustness, verifiability, provenance) in three layers — formal objective, structured semi-objective, human residual — with fatal conditions as a hard gate that soft quality cannot compensate. Coverage splits into specification, test and state coverage; mutation score measures test strength; verification and specification are separate axes (a perfect proof of the wrong theorem solves nothing); evaluator agreement is not truth; cost is not quality unless the specification makes it one. Quality evidence grades run from E (surface validity) to A+ (formal verification plus goal alignment).
- TheoryTHY-2026-0107Binary residual quality measurement (IBQF / BRQM)A 0–10 rating asks the respondent to perceive, build a reference, calibrate a scale, integrate dimensions and map to a number; the burden is highest exactly when the measured state is heaviest. BRQM instead takes many local, concrete, single-construct, non-numeric binary or pairwise answers b_i ∈ {0,1} and lets the measurement system reconstruct a latent multidimensional quality θ̂_H with Bradley–Terry / Thurstone / IRT-type models, adaptive item selection by information gain per human cost, blind and counterbalanced designs, and an explicit rater-disagreement structure — because binary observation is not binary phenomenon and disagreement is not error.
- TheoryTHY-2026-0108Typed, versioned quality ontology for high-ambiguity artifactsBefore any item is written, quality must be a typed space Q[domain, task, context, audience] = Q_core ⊕ Q_domain ⊕ Q_task, with a construct graph of dependencies and conflicts and a measurement itemization pipeline Task → Construct → Indicator → Item → Observation → Latent estimate. Any metric (BLEU, CLIPScore, aesthetic model score) is one projection of the space; a construct validity gate (coverage, discriminant validity, convergent evidence, context stability) guards against measuring the wrong thing precisely; the ontology is open — new constructs may be added from residual errors — but every revision is a version, and multimodal quality is not the mean of modality scores.
- TheoryTHY-2026-0109Scaffolding capability record: SSR, SDR, SCM and the ablation ladderA system is (model M, scaffolding vector S) with S = (tool, retry, multi-sample, verifier, environment, memory, planner). Single-pass quality Q_SP = Q(M, 0) and full quality Q_F = Q(M, S_F) define the scaffolding survival ratio SSR = Q_SP / Q_F, the dependence ratio SDR = 1 − SSR, the scaffold cost multiplier SCM_j = C_j^F / C_j^SP and marginal yields along a controlled ladder A0 native single pass → A1 extended budget → A2 multi-sample → A3 verifier → A4 tool/environment → A5 full agentic. The record also carries scaffold interactions (Shapley-style attribution), discarded and hidden work, and a capability vector (single-pass, loop, tool, verification, environment intelligence). A loop is not cheating; high dependence is a structure, not a defect — what destroys comparability is hiding the capability source and its physical cost.
- TheoryTHY-2026-0110The canonical intelligence event, Pareto comparison and no premature scalarizationIntelligence is not a score of a model but an event 𝔍_IPM = (task object, quality object, semantic work object, physical computation object, scaffolding capability record, measurement metadata); the research object is the relation physical computation → effective semantic work → quality. Given a declared projection Q*, an intelligence yield vector (Q*/E_marg, Q*/V_C, Q*/V_M, Q*/B_M, Q*/B_N, Q*/T_wall) and semantic yields split efficiency into physical→semantic and semantic→outcome stages. Systems are compared on Pareto frontiers — single-pass, scaffolded, and their gap — under the rule 'vector before score, structure before average, uncertainty before false precision'; four capability archetypes (native, efficiently scaffoldable, compute-amplified, environment-coupled) are descriptive, not a ranking. A minimum reporting standard and a grade bundle (Q, μ, E, CST, S) make every claim carry its boundary and uncertainty. IPM is a metrology candidate, not a discovered natural constant.