EVEMISSLAB

ExperimentEXP-2026-0022v0.1

PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)

Sixteen hand-authored natural-language tasks across the preregistered families; A/B/C share one candidate ledger and equal accounted budgets; C receives no gold constraint or supersession metadata; the gold rubric is visible only to a condition-blind judge; literal machine checks are independent of the judge; identical answers reuse one judge cache key; the system fails closed without OPENAI_API_KEY; a frozen-cache replay provider allows exact replay after a live run. 25 tests pass. The execution runtime had no API key, so no real model output exists; the bundled mock smoke file is marked MOCK_ONLY_NOT_REAL_MODEL.

Research status
STABLE the current conclusions are relatively stable
Evidence level
E1 Internal observation
Result
INCONCLUSIVE
Data basis
NOT RUN Designed and validated as a harness; never executed with a real model.
Version
0.1
Updated
2026-09-09
Created
2026-09-09
Domain
Reasoning, Evaluation
Program
PRG-2026-0001 Adaptive Epistemic Systems
Authors
Neo.K (EveMissLab)
AI collaborators
Sol (GPT-5.6, OpenAI ChatGPT)

Hypothesis

hypothesis
A real language model under the PACC runtime shows the coherence / valid-novelty gains and recoverable breadth loss seen in the synthetic witness.

Setup

benchmark_ids
software_environment
Python; OpenAI API provider (fails closed without key); deterministic fake provider for protocol tests only.

Runs

run_count
0
metrics
verdict
REAL_LLM_HARNESS_VALIDATED_BUT_REAL_MODEL_NOT_EXECUTED
execution_status
NOT_EXECUTED_REAL_MODEL
tests_passed
25
tasks
16
real_llm_calls
0

Interpretation

interpretation
Closes the harness, not the scientific question. Next action: a small real smoke (6 tasks × 2 repetitions × 3 candidates), freeze the cache, inspect blind-evaluator consistency, then the full 16-task primary without changing prompts or metrics.

Limitations

limitations
  • No claim about real-model reasoning, intent understanding, imagination, human-rated usefulness, cross-model transfer or hallucination is permitted before a real-model result exists.
  • A single-model judge is not human evaluation even after a live run.
  • Executed for the first time on 2026-09-11 with a local open-weight model — see EXP-2026-0023.

Reproduction

reproduction_instructions
Extract PACC-Hybrid-Lab_v0.2_REAL_LLM_HARNESS_FINAL.zip; python -m pytest -q (25 tests); set OPENAI_API_KEY and run the smoke per docs/REPRODUCIBILITY_v0.2.md.

Recorded fields

random_seeds

    Relations

    SourceRelationTargetStatusID
    EXP-2026-0022 PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)runs_onSYS-2026-0003 PACC-LLM Hybrid LabACTIVEREL-2026-0320
    EXP-2026-0022 PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)uses_benchmarkBEN-2026-0003 PACC-LLM Hybrid A/B/C benchmarkACTIVEREL-2026-0321
    EXP-2026-0022 PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)extendsEXP-2026-0021 PACC-Hybrid v0.1 — synthetic A/B/C architecture witnessACTIVEREL-2026-0322
    EXP-2026-0022 PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)testsTHY-2026-0002 Canonical symbolic state and candidate → verify → commit authorityACTIVEREL-2026-0323
    EXP-2026-0022 PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)producedART-2026-0035 PACC-Hybrid-Lab v0.2 REAL LLM HARNESS FINAL artifact://evemisslab/adaptive-epistemic-systems/PACC-Hybrid-Lab_v0.2_REAL_LLM_HARNESS_FINAL.zipACTIVEREL-2026-0324
    EXP-2026-0023 PACC-Hybrid v0.2 — first real-model run, on a local 9B open-weight modelextendsEXP-2026-0022 PACC-Hybrid v0.2 — real-language-model A/B/C harness (not yet executed)ACTIVEREL-2026-0329

    History and provenance

    Canonical URL
    https://evemisslab.com/ai/experiments/EXP-2026-0022/
    Machine-readable
    /ai/experiments/EXP-2026-0022/index.json
    Snapshot
    AI-SNAPSHOT-v0.1-fe85b9694a45
    Provenance
    source
    EveMissLab research collection: Adaptive Epistemic Systems (真本體論13)
    extracted_by
    Splice (Claude Code), reading the canonical UTF-8 sources and each lab's own result reports
    extracted_at
    2026-09-11
    generator
    tools/extract_aes/extract.py
    claim_boundary
    status, evidence level and result type follow the source artifact's own stated claim boundary; nothing is upgraded beyond what the report supports