Main result
For the small state-only controller, erasing executable abstract state changed half of the raw decisions, while changing an irrelevant placebo did not. This supports the narrow claim that the declared state affected behavior in this synthetic control. It does not show open-ended coding ability, and public verification can hide weak raw proposals.
- Keep: typed state slots, explicit legality, and separate raw/full evaluation.
- Test next: tasks where state is necessary but not sufficient, with independently specified labels.
- Do not promote: general Python or repository coding from generated microtasks.
Corrections
The two-parameter predicate gate performs routing from a supplied bit; it does not infer semantics. The semantic_rule_gate and repository_bundle_gate placebo controls are tautological because the nuisance variable is absent from the gate, so their historical causal rates cannot support the architecture. Earlier latency measurements started after inference and were not end to end. New runs report inference, selection, scoring, and inference-through-selection separately.
Experiment 14 is a CPU-only smoke checkpoint and is paused before its planned multi-seed run. The actual Raly compiler/runtime is not in the overnight execution path: this is Raly-style or Raly-inspired research.
Planned benchmark sequence
HumanEval+ will be used as a development benchmark, so task-level failures may be inspected and optimized against. Any tuned score will be reported as development performance, not held-out evidence. MBPP+ is the planned cross-benchmark check, followed by BigCodeBench-Hard Complete and a later time-separated LiveCodeBench sample.
Autoresearch uses separate frozen local proxy tasks. Public prompts and solutions stay out of training; HumanEval+ is the deliberate, disclosed exception for iterative diagnosis. Future runs preregister <=9M parameters, raw/full/null scores, search/test-time budget, and separate latency fields. No public benchmark was run here.
Methods and data
The repository contains the full dashboard, append-only JSONL records, invalidated oracle log, preregistrations, generator audits, and reproduction commands.
Experiment 13 findings →
Full dashboard source →
Back to the project →