Results — bulk-matter track (D0–D4)

Worlds: modular tanh lattices — collective bistability, hysteresis, per-region timescales, slow fatigue (D4). One collective switch per region; the agent must calibrate scrambled sensors, find the switches, map thresholds and memory. World guide.

1. Difficulty tracks ability (chat tier, M0 grid)

model (null harness)D0D1D2D3
google/gemini-3.5-flash0.720.580.540.35
deepseek/deepseek-v4-flash0.640.560.520.38
openai/gpt-5-nano0.120.110.130.04
scripted reference0.720.720.650.67

Pooled difficulty↔reward Spearman ρ=−0.50 (p=0.003, n=32 rollouts). gpt-5-nano sits at the null floor everywhere — mostly protocol failure, which motivated the coding-harness tier.

2. The D4 frontier (tools tier, 5 seeds)

pairingpredictionbudget used
claude-fable-5 + claude_code0.349 ± 0.040.32
claude-opus-5 + claude_code0.324 ± 0.110.41
gpt-5.6-sol + codex0.244 ± 0.070.54
gemini-3.1-pro + codex0.194 ± 0.120.01
scripted reference / tail / null0.274 / 0.211 / 0.1330.11 / — / —

3. Preparation and theory (M2/M3 grid, D4, 3 seeds, n_prep=3)

pairingpredictionpreparationtheory
claude-fable-50.3050.730.29 (3/3 submitted)
claude-opus-50.2690.670.31 (2/3)
gpt-5.20.2190.470/3 submitted
gpt-5.6-sol0.2140.510.21 (1/3)
prep_pi certifier / null0.34 / 0.1450.24 / 0.0

4. Best artifacts

5. Cost accounting (tools tier, D4)

pairingmodel callscached promptcompletionwall time
gpt-5.6-sol + codex66–984–5M12–17k8–12 min
claude-opus-5 + claude_code67–934–6.5M111–125k25–36 min
claude-fable-5 + claude_code113–17322–51M191–241k~1 h

fable's edge rides a 5–10× cached-context bill and hour-long rollouts — a long-context research strategy the others don't attempt.

Generated by python -m physim.results; numbers from REPORT.md (the running lab log with full methodology).