Measuring evolved worlds: a closed-book science pilot

Post 11 · Current iteration: BLOB2v2r2 · Updated after the 2026-09-06 budget stop and completed-run audit. Post 10's archive supplies the worlds; this post asks what a coding agent can learn through instruments.

Current status: stopped, not a completed benchmark. The resource revision is implemented. The user canceled further queued work; all evaluation processes are now stopped. No future benchmark is authorized. The two reported completions below are diagnostic pilot evidence, not a trusted comparative benchmark: both audits found non-composing replies during parallel exploration. No live state-integrity fix was applied.
BLOB2v2r2 taskReported outcomeWall time
E1 #928Completed: 0.680662383 h 34 m 48 s
E2 #942Completed: 0.654498445 h 49 m 22 s
E1 #929HarnessError, provider/server mid-response failure; unscored
E2 #943Entered startup before the operator fence; canceled, no retained science result~11 min
E1 #930Canceled during setup, before the model stage
E2 #944Not admitted

These are Fable-5 pilot outcomes, not cohort means or a model ranking. Both completed cases now have process audits. Failures and cancellations are not scored zeros. See the scope and terminal stop record.

1 · The current task: known syllabus, hidden instances

The agent studies an unknown spatial dynamical system through two remote sensor clusters, with 13 and 19 slots. Ports are anonymous; sensor geometry, actuator meanings, and emission coupling are undisclosed. Per-port global mean/variance is available. The evaluator's worlds are p4g2_044 (E1, sparse stable blobs) and p6g8_033 (E2, reorganizing labyrinth and rotors).

  1. Explore. The agent knows the SYLLABUS: categories, payload shapes, scoring references, and sampling domains, not the drawn instances. It can read the entire base record, t=0–2500 tu in 5-tu steps, and fork from any base grid time or an existing fork. Forks are specified to have independent noise streams. It can read, wait, adjust, inject, and discard forks. Its experiments are not themselves scored.
  2. Choose ready. probe_ready() is agent-triggered and irreversible. It reveals all six concrete instances and closes every world tool, including base reads. There is no automatic reveal at a span end.
  3. Answer closed-book. Own notes, recorded data, code, and offline computation remain available, including code edits. Only probe_status and probe_submit remain among world interface tools. Last accepted submissions win.

Anchors are drawn from [600, 2300], at simulation-step resolution rather than the 5-tu read grid. The shipped syllabus specifies six instances per episode:

CategoryPrediction
L1Device-0 readings after a hidden accepted actuator sequence
L2Readings of an additional hidden 13-slot sensor cluster
L3FUndisturbed streams at 5/25/100/400-tu horizons; E2 omits 5
L3E / L3SE1: crossing counts in sixteen 50-tu windows. E2: global mean/variance in 200-tu windows ending at anchor+400/+800
L4 / L4DEmission response at revealed port, amplitude, duration, and lags: amplitude 1.5–3.0 / 0.30–0.90 respectively
skill = clip(1 − CRPS_agent / CRPS_best_reference, −1, +1)
reward = mean skill over six instances; unsubmitted = −1

Each answer supplies a mean and predictive standard deviation. References use pre-anchor history: climatology/persistence, AR(2) for L3F, and event-rate references for L3E. Zero means matching the reference, not recovering a law. Agents can collect later recorded history before reveal; the reference does not have that information. This distinction matters below.

2 · Resource policy: safety guards, not science budgets

BLOB2v2r2-E1/E2 changes resource handling, not worlds, seeds, instance draws, truth arrays, or scoring. The seven private guards are:

GuardCeiling
Cumulative fork spawns / logically open handles100,000 / 10,000
Aggregate forward simulation / sensor meter10,000,000 tu / 1,000,000,000 node-tu
Actuator / emission meters10,000,000 commanded cu / 1,000,000 emission-price units
Persisted operation-log plus emission entries1,000,000

These are high runaway guards, not promises of unlimited capacity. Agents see no budget, remaining allowance, or efficiency reward. There is no per-experiment or full-rollout wall-time deadline in this policy. The apparatus still limits emission amplitude to 1.0 and duration to 50 tu; those bounds do not limit experiment length. A guard trip terminates the run as an unscored resource failure, not a forced reveal or science score.

An eight-entry process-wide LRU cache limits resident live states, not logical forks. Eviction does not close a handle. A later cache miss reconstructs its state from the fork stream and operation log. Reconstruction is not charged as logical forward time. Eight states is not a total-memory bound, and response-size guards are transport limits, not experiment deadlines.

3 · Two completed cases: E1 #928 and E2 #942

Both audits follow one exact trace each: E1 #928 (0bdd699154ee4e1d96aac4e0961bc11d) and E2 #942 (ae982494a72144c186f58a687a99cd33). The paired notes hold every exact tool ID and artifact line cited here. The two cases use different worlds, menus, ports, and hidden draws. Their scores are not averaged, and their difference is not evidence about model quality.

QuantityE1 #928E2 #942
World / menup4g2_044 · L3Ep6g8_033 · L3S
Reported episode skill (six-instance mean)0.680662380.65449844
Wall time3 h 34 m 48 s5 h 49 m 22 s
Completed model calls / error attempts1,216 / 232,218 / 40
Prompt tokens (excludes cache reads, includes cache writes)4,818,5505,979,370
Cache-read tokens117,000,774213,938,314
Output tokens (reasoning subset)1,199,851 (153,020)2,265,077 (97,241)
Input + output tokens123,019,175222,182,761
Agent delegations / Bash / Write22 / 324 / 5235 / 665 / 276
Environment requests / persisted turns1,065 / 8911,300 / 1,023
Environment timeouts / missing results54 / 2266 / 11
Live simulated tu / fork read frames11,495 / 62514,655 / 895
Fork replies / distinct returned IDs181 / 145100 / 96
Peak logical open / resident forks31 / 812 / 8
Cap hits / resource truncation0 / none0 / none
Post-ready world requests / accepted submissions0 / 60 / 6

Token counts follow the installed verifiers Usage schema. Error attempts have no recorded usage; auxiliary token-count requests are not recorded; separate cache-write totals and billed dollars are unavailable. Reasoning tokens are already inside output tokens. Wall time is not a billing measurement. Both runs stayed far below every private guard.

What each agent actually submitted

InstanceE1 #928 method → skillE2 #942 method → skill
L1Recorded device readings contracted toward global means with fitted actuator features → 0.2225Fitted actuator-distance contraction toward a local mean → 0.5225
L2Global nowcast/climatology repeated across hidden slots → −0.0034Post-ready local blend: 0.7 pooled-device mean + 0.3 global mean, repeated across slots → 0.1260
L3FCatmull–Rom record interpolation; all targets inside base 2500 → 0.9964Cubic record interpolation; targets 1814.02 / 1889.02 / 2189.02, inside base → 0.9564
L3E / L3SCrossing counts from interpolated record through 2494.78 → 0.8856Window means of saved global streams; second window 2660.72–2860.72 is outside base but covered by 16 global samples collected before ready → 0.9518
L4Undisturbed recorded mean; response envelope in sigma; port 3 → 0.9891Measured-RMS shrinkage of mean and sigma; port 6, amp 2.7998 → 0.4040
L4DUndisturbed recorded mean; envelope in sigma; port 1 → 0.9938Same RMS/shrinkage model; port 11, amp 0.8124 → 0.9664

Neither submitted path contains a general learned dynamical simulator. That is not the same as “no modeling”: E2's actuator contraction and RMS→decorrelation/shrinkage models are genuine, narrower empirical models. E1's two drawn dose means stayed passive; its fitted port-2 mean template was unused because the drawn ports were 3 and 1. E2's emission model changes the mean itself, backed by measured late RMS about 0.22 for port 6 versus about 0.01 for port 11 at its training condition. E2's L4 draw (amplitude 2.7998, above the apparatus limit 1) scored 0.404 with bounded heuristic scaling and no direct training validation above amplitude 1.

Record and continuation lookup were permitted. Both agents used the readable base record heavily. E2 also sampled a fresh anchor-2500 continuation before ready: 16 global observations at 2525–2900 in 25-tu steps, which cover its out-of-base L3S window. Its device continuation is sparser and has a 2800→2905 gap; it is not uniform 5-tu coverage. High undisturbed-forecast scores therefore show careful data collection and interpolation, not learned extrapolation. Both L2 answers repeat marginal estimates across hidden slots; neither recovers hidden sensor geometry.

Post-ready behavior was clean in both cases: zero sampled world-tool calls after ready; only local file, predictor, and submission work; all six payloads exactly match the archived payload files. E1 submitted L3E, L1, L2, L3F, L4, L4D; E2 submitted L3S, L1, L2, L3F, L4D, L4.

4 · State integrity limits the interpretation

In E1, distinct parallel fork requests for anchors 200 and 600 returned the same ID at different anchors; another batch repeated this across four anchors, and parallel advancing reads returned repeated times. In E2, the pattern is weaker but present: 100 successful fork replies returned 96 distinct handles, one handle was reported at anchors 2500 and 1000, a wait reply reported 1250 followed by a zero-window read at 1225, and five handles acknowledged reset twice (E2 evidence). Native transport performs whole-state GET → tool → PUT with last-write-wins updates. This strongly supports lost updates, not random 128-bit ID collisions. Transaction logs are missing, so exact interleavings cannot be reconstructed.

Reported counters are not exact physical-work ledgers. E1's ordinary replies exceed persisted turns by 118 and its reply sums exceed the meters; E2's ordinary replies equal its 1,023 persisted turns, yet reply-derived sensor and simulation sums fall below the final meters. A matching count is not transaction exactness. E2 also had 266 environment timeouts; a timeout does not tell us whether server work committed. Completion and six accepted submissions do not repair this evidence gap. No live fix was applied.

5 · Before any further model comparison

Proposed test ladder only; no new runs or fixes are claimed here.

  1. Deterministic native checks first: transport/concurrency, phase closure, serialization, replay, resets, meters, and artifact persistence. Use the actual shared-state path. Add record-interpolation and no-emission controls to separate high scores from learned dynamics.
  2. One short, cheap-model wiring smoke: check tools, files, arguments, reveal, and six submissions. Label it diagnostic, not scientific performance.
  3. One full frontier pilot: allow normal investigation, then review state integrity, artifacts, methods, failures, and cost before expanding.
  4. Only then a predeclared small matched seed/world/model panel: separate within-record from out-of-record prediction and weak from active emission responses. Report failures; do not select favorable cases afterward.

A weak model's success or failure is not a judgment of scientific difficulty. A short smoke can miss long-context and parallel-state failures. Cheap wiring checks cannot replace native controls or a full pilot audit.

History, kept separate

The historical BLOB2-E1/E2 round-4 results belong to the earlier budgeted, fixed-span contract system. The subsequent capped BLOB2v2 cohort remains separate, as do invalid-mode attempts. None are pooled into BLOB2v2r2. Earlier interface and curriculum lessons remain useful; their scores are not a cross-version improvement claim.

Series: index · previous: Evolving at scale · next: Breeding spatial economies