| BLOB2v2r2 task | Reported outcome | Wall time |
|---|---|---|
| E1 #928 | Completed: 0.68066238 | 3 h 34 m 48 s |
| E2 #942 | Completed: 0.65449844 | 5 h 49 m 22 s |
| E1 #929 | HarnessError, provider/server mid-response failure; unscored | — |
| E2 #943 | Entered startup before the operator fence; canceled, no retained science result | ~11 min |
| E1 #930 | Canceled during setup, before the model stage | — |
| E2 #944 | Not admitted | — |
These are Fable-5 pilot outcomes, not cohort means or a model ranking. Both completed cases now have process audits. Failures and cancellations are not scored zeros. See the scope and terminal stop record.
The agent studies an unknown spatial dynamical system through two remote
sensor clusters, with 13 and 19 slots. Ports are anonymous; sensor geometry,
actuator meanings, and emission coupling are undisclosed. Per-port global
mean/variance is available. The evaluator's worlds are p4g2_044
(E1, sparse stable blobs) and p6g8_033 (E2, reorganizing labyrinth
and rotors).
SYLLABUS: categories,
payload shapes, scoring references, and sampling domains, not the drawn
instances. It can read the entire base record, t=0–2500 tu in 5-tu steps,
and fork from any base grid time or an existing fork. Forks are specified to have independent
noise streams. It can read, wait, adjust, inject, and discard forks. Its
experiments are not themselves scored.probe_ready() is agent-triggered and
irreversible. It reveals all six concrete instances and closes every world
tool, including base reads. There is no automatic reveal at a span end.probe_status and probe_submit remain among world
interface tools. Last accepted submissions win.Anchors are drawn from [600, 2300], at simulation-step resolution rather than the 5-tu read grid. The shipped syllabus specifies six instances per episode:
| Category | Prediction |
|---|---|
| L1 | Device-0 readings after a hidden accepted actuator sequence |
| L2 | Readings of an additional hidden 13-slot sensor cluster |
| L3F | Undisturbed streams at 5/25/100/400-tu horizons; E2 omits 5 |
| L3E / L3S | E1: crossing counts in sixteen 50-tu windows. E2: global mean/variance in 200-tu windows ending at anchor+400/+800 |
| L4 / L4D | Emission response at revealed port, amplitude, duration, and lags: amplitude 1.5–3.0 / 0.30–0.90 respectively |
Each answer supplies a mean and predictive standard deviation. References use pre-anchor history: climatology/persistence, AR(2) for L3F, and event-rate references for L3E. Zero means matching the reference, not recovering a law. Agents can collect later recorded history before reveal; the reference does not have that information. This distinction matters below.
BLOB2v2r2-E1/E2 changes resource handling, not worlds, seeds,
instance draws, truth arrays, or scoring. The
seven private guards are:
| Guard | Ceiling |
|---|---|
| Cumulative fork spawns / logically open handles | 100,000 / 10,000 |
| Aggregate forward simulation / sensor meter | 10,000,000 tu / 1,000,000,000 node-tu |
| Actuator / emission meters | 10,000,000 commanded cu / 1,000,000 emission-price units |
| Persisted operation-log plus emission entries | 1,000,000 |
These are high runaway guards, not promises of unlimited capacity. Agents see no budget, remaining allowance, or efficiency reward. There is no per-experiment or full-rollout wall-time deadline in this policy. The apparatus still limits emission amplitude to 1.0 and duration to 50 tu; those bounds do not limit experiment length. A guard trip terminates the run as an unscored resource failure, not a forced reveal or science score.
An eight-entry process-wide LRU cache limits resident live states, not logical forks. Eviction does not close a handle. A later cache miss reconstructs its state from the fork stream and operation log. Reconstruction is not charged as logical forward time. Eight states is not a total-memory bound, and response-size guards are transport limits, not experiment deadlines.
Both audits follow one exact trace each:
E1 #928
(0bdd699154ee4e1d96aac4e0961bc11d) and
E2 #942
(ae982494a72144c186f58a687a99cd33). The
paired notes
hold every exact tool ID and artifact line cited here. The two cases use
different worlds, menus, ports, and hidden draws. Their scores are not averaged,
and their difference is not evidence about model quality.
| Quantity | E1 #928 | E2 #942 |
|---|---|---|
| World / menu | p4g2_044 · L3E | p6g8_033 · L3S |
| Reported episode skill (six-instance mean) | 0.68066238 | 0.65449844 |
| Wall time | 3 h 34 m 48 s | 5 h 49 m 22 s |
| Completed model calls / error attempts | 1,216 / 23 | 2,218 / 40 |
| Prompt tokens (excludes cache reads, includes cache writes) | 4,818,550 | 5,979,370 |
| Cache-read tokens | 117,000,774 | 213,938,314 |
| Output tokens (reasoning subset) | 1,199,851 (153,020) | 2,265,077 (97,241) |
| Input + output tokens | 123,019,175 | 222,182,761 |
| Agent delegations / Bash / Write | 22 / 324 / 52 | 35 / 665 / 276 |
| Environment requests / persisted turns | 1,065 / 891 | 1,300 / 1,023 |
| Environment timeouts / missing results | 54 / 2 | 266 / 11 |
| Live simulated tu / fork read frames | 11,495 / 625 | 14,655 / 895 |
| Fork replies / distinct returned IDs | 181 / 145 | 100 / 96 |
| Peak logical open / resident forks | 31 / 8 | 12 / 8 |
| Cap hits / resource truncation | 0 / none | 0 / none |
| Post-ready world requests / accepted submissions | 0 / 6 | 0 / 6 |
Token counts follow the installed verifiers Usage schema.
Error attempts have no recorded usage; auxiliary token-count requests are not
recorded; separate cache-write totals and billed dollars are unavailable.
Reasoning tokens are already inside output tokens. Wall time is not a billing
measurement. Both runs stayed far below every private guard.
| Instance | E1 #928 method → skill | E2 #942 method → skill |
|---|---|---|
| L1 | Recorded device readings contracted toward global means with fitted actuator features → 0.2225 | Fitted actuator-distance contraction toward a local mean → 0.5225 |
| L2 | Global nowcast/climatology repeated across hidden slots → −0.0034 | Post-ready local blend: 0.7 pooled-device mean + 0.3 global mean, repeated across slots → 0.1260 |
| L3F | Catmull–Rom record interpolation; all targets inside base 2500 → 0.9964 | Cubic record interpolation; targets 1814.02 / 1889.02 / 2189.02, inside base → 0.9564 |
| L3E / L3S | Crossing counts from interpolated record through 2494.78 → 0.8856 | Window means of saved global streams; second window 2660.72–2860.72 is outside base but covered by 16 global samples collected before ready → 0.9518 |
| L4 | Undisturbed recorded mean; response envelope in sigma; port 3 → 0.9891 | Measured-RMS shrinkage of mean and sigma; port 6, amp 2.7998 → 0.4040 |
| L4D | Undisturbed recorded mean; envelope in sigma; port 1 → 0.9938 | Same RMS/shrinkage model; port 11, amp 0.8124 → 0.9664 |
Neither submitted path contains a general learned dynamical simulator. That is not the same as “no modeling”: E2's actuator contraction and RMS→decorrelation/shrinkage models are genuine, narrower empirical models. E1's two drawn dose means stayed passive; its fitted port-2 mean template was unused because the drawn ports were 3 and 1. E2's emission model changes the mean itself, backed by measured late RMS about 0.22 for port 6 versus about 0.01 for port 11 at its training condition. E2's L4 draw (amplitude 2.7998, above the apparatus limit 1) scored 0.404 with bounded heuristic scaling and no direct training validation above amplitude 1.
Record and continuation lookup were permitted. Both agents used the readable base record heavily. E2 also sampled a fresh anchor-2500 continuation before ready: 16 global observations at 2525–2900 in 25-tu steps, which cover its out-of-base L3S window. Its device continuation is sparser and has a 2800→2905 gap; it is not uniform 5-tu coverage. High undisturbed-forecast scores therefore show careful data collection and interpolation, not learned extrapolation. Both L2 answers repeat marginal estimates across hidden slots; neither recovers hidden sensor geometry.
Post-ready behavior was clean in both cases: zero sampled world-tool calls after ready; only local file, predictor, and submission work; all six payloads exactly match the archived payload files. E1 submitted L3E, L1, L2, L3F, L4, L4D; E2 submitted L3S, L1, L2, L3F, L4D, L4.
In E1, distinct parallel fork requests for anchors 200 and 600 returned the same ID at different anchors; another batch repeated this across four anchors, and parallel advancing reads returned repeated times. In E2, the pattern is weaker but present: 100 successful fork replies returned 96 distinct handles, one handle was reported at anchors 2500 and 1000, a wait reply reported 1250 followed by a zero-window read at 1225, and five handles acknowledged reset twice (E2 evidence). Native transport performs whole-state GET → tool → PUT with last-write-wins updates. This strongly supports lost updates, not random 128-bit ID collisions. Transaction logs are missing, so exact interleavings cannot be reconstructed.
Reported counters are not exact physical-work ledgers. E1's ordinary replies exceed persisted turns by 118 and its reply sums exceed the meters; E2's ordinary replies equal its 1,023 persisted turns, yet reply-derived sensor and simulation sums fall below the final meters. A matching count is not transaction exactness. E2 also had 266 environment timeouts; a timeout does not tell us whether server work committed. Completion and six accepted submissions do not repair this evidence gap. No live fix was applied.
Proposed test ladder only; no new runs or fixes are claimed here.
A weak model's success or failure is not a judgment of scientific difficulty. A short smoke can miss long-context and parallel-state failures. Cheap wiring checks cannot replace native controls or a full pilot audit.
The historical BLOB2-E1/E2 round-4 results belong to the earlier budgeted, fixed-span contract system. The subsequent capped BLOB2v2 cohort remains separate, as do invalid-mode attempts. None are pooled into BLOB2v2r2. Earlier interface and curriculum lessons remain useful; their scores are not a cross-version improvement claim.