Eunoia Omega
Reasoning that should get cheaper the more of it you do — at a fraction of the scale the field currently assumes is necessary.
- Parameter budget
- ≤ 450M
- Language pretraining
- None
- Domain
- Abstract & interactive reasoning
- Evaluation
- ARC-AGI 1, 2 and 3
- Cost target
- Under $10 per run
- Availability
- Not released
Every competitive system starts each problem from zero.
The systems currently leading interactive reasoning benchmarks pair a frontier language model with an external harness. They work — but their priors about how environments behave come from web-scale pretraining, and their weights are frozen. Published evaluation runs cost between $400 and $2,986.
The consequence is the part we find scientifically decisive: a frozen system rediscovers that obstacles block movement on its fortieth environment exactly as expensively as on its first. Cost per environment does not fall with experience. Human learners do not behave this way.
Capability comes from the form of the representation, not the size of the model that produces it.
If that holds, the representation can be built directly and at small scale, without first acquiring a trillion parameters of unrelated knowledge — and it can be made to improve with experience, which a frozen prior cannot.
Omega is the attempt to test that proposition rather than argue it.
The claim, stated before the model existed.
Cost per environment should decrease as environments accumulate. Frozen-weight systems should stay flat. That difference is measurable on the metric the benchmark already defines, it is independent of leaderboard position, and it cannot be bought with scale.
We wrote the evaluation protocol, the seed discipline, and the abandonment threshold for every component before construction began. If the transfer curve is flat across five seeds, the claim is unsupported and we will publish it as unsupported.
- No configuration comparison is drawn from a single run.
- Every reported score carries seed-level dispersion, not a point estimate.
- An effect smaller than seed dispersion is not an effect and is not reported.
- The protocol is versioned and is not modified mid-programme.
What we are not claiming
- No trained results exist. Nothing here is a statement about achieved performance.
- We are not aware of any published from-scratch attempt at this benchmark in this parameter regime, so our own score targets are hypotheses without calibration.
- A null result on the primary claim is a publishable outcome and will be reported as one.
What we publish, and what we keep.
Released
- Evaluation protocol and its version history
- Benchmark results, including negative ones
- Training curves, hardware and compute
- Ablations and error analysis
Retained
- Corpus contents and generator internals
- Cleaning, deduplication and admission pipelines
- Representation design and program space
- Preprocessing code and source manifests
This boundary is fixed in advance and enforced by provenance records, not by discretion.


