Can AI agents do science? Procedurally generated universes with hidden laws behind anonymous interfaces โ agents explore, experiment, build instruments, and are scored on prediction, state preparation, and executable theories.
Headline findings and leaderboards, one page per world track: bulk matter (D) ยท chemistry (C) ยท ecology (B) ยท evolution (E).
How an AI scientist is graded: prediction contracts scored by a proper distributional rule (CRPS vs god-mode truth ensembles, noise-floor subtracted), preparation policies run on fresh clones, executable theories replayed against the laws โ plus the reference ladder (null โ tail โ persistence โ compact oracle โ replication) and the measurement-adequacy audit that certifies the contracts can see each world's emergent phenomena.
A research program building the next world from scratch: three coupled fields whose excitations are particle-like blobs that move (drift bifurcation), bind into molecules (d*=15.70), come in port-distinguishable species, and self-assemble into a working machine โ a relay tug hauling cargo upstream against a load field at 5.8ร baseline; then evolved at scale (423-cell archive) and turned into a measurement benchmark for AI scientists (post 11). Full report with films, equations, the dial table, and honest negatives.
Plain-language guide to the simulated universes, one page per track with god-view films of the dynamics (waves entraining, ecosystems starving, gene pools evolving through storm cycles): bulk matter ยท chemistry ยท ecology ยท evolution.
What the agents actually did: narrative experiment logs with timeline figures, the files each agent wrote (its instruments and theories), and contract answers vs ground truth โ split by world track.
The verifiers taskset, engine, and docs generators. Raw rollout traces: HF dataset. Full running lab log: REPORT.md.