physim asks: can AI agents do science? Each task is a procedurally generated world with hidden laws behind an anonymous port interface. Agents explore under a tick budget, then are scored on prediction (held-out protocols on fresh world copies), preparation (submit control policies that steer the world into target states), and optionally an executable theory (a simulator of the sensors, replayed against ground truth). Since v0.10.0 prediction is scored distributionally (CRPS against god-mode truth ensembles — honest uncertainty and multimodal structure pay; how scoring works). Four world tracks: bulk matter (D0–D4), chemistry (C0–C4), ecology (B0–B2), evolution (E0–E1).
| claim | evidence |
|---|---|
| Difficulty tracks model ability | reward falls monotonically D0→D4 for every competent model; pooled Spearman ρ=−0.50 (p=0.003) on the first grid |
| The frontier is unsaturated on hard worlds | best D4 pairing (claude-fable-5 + claude_code): 0.35 prediction vs 0.93 replication ceiling; S1 (slow-adaptation stratum) ≤0.5 for all models |
| Preparation ≫ prediction at the frontier | on D4, fable-5 steers worlds at 0.73 success while predicting them at 0.305 — control forgives model error; scripted PI baseline manages 0.24 |
| Executable theories track understanding | theory score correlates with prediction accuracy (r=0.72); best D2 theory scored 0.895, best C1 theory 0.96 |
| Instrument discovery is measurable — and unsolved | C1's apparatus-forced preparation (a band only reachable by finding and driving the hidden sensor-stage port) scores 0/1 for the best agent while ordinary preparations score 4/4; fable-5 even noticed the gain-apparatus channel behaving oddly ("integrator") without discovering sensor motion |
| Agents rediscover real structure in alien coordinates | claude-opus-5 on chemistry world C1: "cells" (=objects), "line attractors" (=growth), "trigger → absorbing state" (=death, thresholds measured), "immune above σ≈850" (=too big to kill); held-out validation mean |err| 0.021 — read it in the gallery |
| Distributional literacy is a conduct trait | given optional quantile answers under a proper score (CRPS), fable used them on 5/5 stochastic-world rollouts, gpt-5.2 on 0/5 — same tools, same prompt; the reward now prices honest uncertainty and gpt leaves it on the table (e.g. E1: fable 0.72 with quantiles vs its own 0.57 legacy point ceiling) |
| Harness matters; models differ in conduct, not just skill | coding harness doubles chat-tier scores (budget → offline fits); gemini-3.1-pro satisfices (~12 turns, 1% budget) where fable runs 60–180-turn research programs |
Under CRPS (2026-02-16 re-runs, distributional oracles): C4 (0.5–0.6) > B2 (0.45–0.63) > B0 (0.13–0.31) ≈ E1 (0.22–0.43) > D4 (0.3–0.45, point-era) > C2/C3 (0.1–0.3) > B1/C0/C1 (≈0). B2's distributional oracle hit 0.975 — its ensembles are wide-but-knowable, and the frontier is nowhere near that.
Scoring upgrade (2026-02-16, v0.10.0): accuracy is now a proper distributional score (CRPS vs the truth ensemble, noise-floor subtracted; answers may include quantiles). On quasi-deterministic worlds (C/D tracks) it reduces exactly to the old point score, so those numbers carry over. B0/B2/E1 frontier rows were re-run under CRPS on 2026-02-16. Details: scoring page.
| pairing | prediction | preparation* |
|---|---|---|
| claude-fable-5 + claude_code | 0.349 ± 0.04 | 0.73 |
| claude-opus-5 + claude_code | 0.324 ± 0.11 | 0.67 |
| gpt-5.6-sol + codex | 0.244 ± 0.07 | 0.51 |
| gemini-3.1-pro + codex | 0.194 ± 0.12 | 0.47† |
| scripted reference / null | 0.274 / 0.133 | 0.24 / 0.0 |
| replication ceiling | ~0.93 | — |
| pairing | track | prediction | preparation | theory |
|---|---|---|---|---|
| gpt-5.6-sol + codex | C1 | 0.97, 0.92 | 1.0 | 0.82 |
| claude-opus-5 + claude_code | C1 | 0.96, 0.95 | 1.0 | 0.96, 0.95 |
| claude-fable-5 + claude_code | C0 | 0.95, 0.91 | 1.0 | 0.96, 0.92 |
| gpt-5.2 + codex | C0 | 0.82, 0.54 | 1.0, 0.75 | — |
| gpt-5.2 + codex (C2, tight scales) | C2 | 0.82, 0.73 | —* | — |
| tail / null baselines (C0) | 0.75 / 0.32 | — | — | |
| tail / null / persistence (C2) | 0.49 / 0.11 / 0.39 | — | — |