B-track worlds are ecosystems (organisms, resources,
competition) built on the same lattice-field engine as chemistry. Agents see
only anonymous ports and sensors; that two kinds of life exist is itself a
discovery. World guide ·
scoring ·
evolution track results →
B0 — the biology track opens (2026-02-14, v0.7.0)
An ecosystem: two organism variants (greedy-fast vs
frugal-fragile) compete for one regenerating resource. Ports fertilize or
poison regions; sensors read species-blind organism density. Emergent laws:
carrying capacity; coexistence under a consumption trade-off; and
natural selection — sustained scarcity drives the greedy variant
extinct while the frugal one inherits the world (verified through the port
interface: poison 32→0). Certified 6/6 seeds.
Legacy point-scored run (2026-02-14): fable 0.63/0.46, gpt-5.2
0.58/0.46, oracle 0.744, floors 0.27/0.18. Same-answer legacy scores in the
re-run (fable 0.39/0.55, gpt 0.49/0.41) confirm skill is unchanged — CRPS
lifts honest answers by not charging the world's own noise floor.
Layer-3 ontology gap: no agent workspace contains ANY biological
vocabulary (population/organism/resource/extinction) — models fit channel
responses to drive levels and miss recovery + extinction strata (S3/S4
0.34–0.55). Third measured rung of the closure ladder (C1 apparatus, C3
species, B0 populations).
Deepest persistence floor of any world (0.15): populations never sit
still.
B1 — selection-boundary worlds (2026-02-14, v0.7.2) and the smoothness ceiling
Richness is alienized across the exclusion threshold:
each instance lands coexist-side or excluded-side, and long boundary-crossing
contracts (S5) probe which. Result: frontier models score 0.77–0.82
(gpt-5.2: 0.77 from just 10 experiments), matching the compact oracle (0.73)
— because population observables respond smoothly to drives, a few
tilt levels interpolate everything. Even near-extinction recovers via
refugia recolonization (no hysteresis at the readout level).
Lesson (recorded as design law): richness is necessary but not
sufficient for frontier-hardness — the observable map must also be
non-interpolable (timing/phase structure, discontinuities, path
dependence), which is exactly why C4's waves break the frontier while B1's
populations don't. Ecology at population granularity is intrinsically
smooth → mid-tier.
Curriculum shipped: B0a (carrying capacity alone) → B0b (pure
competition) → B0 (selection) → B1 (boundary) — each certified separately;
the B-track decomposition ladder is ready for training experiments.
B2 — the ecowave hybrid (2026-02-14): hardness by composition
Waves feed the ecology: wave passage boosts resource
regeneration ("rain"); without waves the world starves. Population tracks
wave rate — including the refractory trap: pacing the medium too fast causes
conduction block and fewer meals (a non-monotonic response that
breaks interpolation). B2 = C4 (waves) + B0a (carrying capacity), composed —
each component separately certified.
agent / reference (CRPS, 2026-02-16 re-run)
prediction
answers
claude-fable-5 + claude_code
0.53, 0.42
quantiles
gpt-5.2 + codex
0.37, 0.34
points
compact hybrid oracle (~55 lines + honest spread)
0.975 (0.92 point-only)
quantiles
tail / null
0.42 / 0.32
points
Legacy point-scored run (2026-02-14): fable 0.48/0.36, gpt-5.2
0.42/0.29, oracle 0.795, floors 0.35/0.28. The distributional oracle at 0.975
shows B2's truth ensembles are wide-but-knowable: an agent that pinned the
laws AND reported honest spread would nearly close the gap — the frontier
(≤0.53) is far from that.
fable's notes contain "wave, pulse, period, rain" — it perceives
the coupling — yet the pacing stratum (S2) scores 0.04–0.31 for every model:
nobody builds the pacing → food-delivery → carrying-capacity chain.
The demonstrated recipe: compose a non-interpolable layer (waves)
with a smooth one (ecology) and the frontier gap returns without scaling.
Hardness lives in coupling structure, not size.
What's next on the ecology track
B3 candidate: two variants + waves — selection whose winner depends on
the wave regime;