Theory and experiments

Mind the Seriality Gap

What denoising steps can compute,
and what our video experiments show.

The three experiments in order, the physics simulator and the score-error audit on side paths, and the argument off the path 3.1 The physics simulator 3.5 Score-error audit 01 The argument 02 Something naive 03 State diffusion 04 State-space guidance The three experiments in order, the physics simulator and the score-error audit on side paths, and the argument off the path 3.1 Physics simulator 3.5 Score-error audit 01 The argument 02 Something naive 03 State diffusion 04 State-space guidance

To-do & ideas ↓ — off the path: what we might run next.

01 / The argument

Why the denoising loop does not provide serial computation — allegedly

CollapseExpand

For us to say that the denoising loop of \(T\) steps provides “serial computation,” we would need to show that allowing \(T\) to grow enlarges the class of problems the model can solve beyond what constant \(T\) can solve. Here, “constant” means independent of input size, with the same limits on the computation in each score evaluation.

SSH’s claim is that, under its assumptions, including the low-score-error condition at a fixed step budget, the tasks in question can already be solved with some constant \(T^*\). Allowing \(T\) to grow then adds no new class of solvable problems. This is the sense in which “the denoising loop does not provide scalable serial computation.” SSH refers to The Serial Scaling Hypothesis.

The easiest case: a perfect score

Let \(c\) be the conditioning frames and \(y^*=f(c)\) the unique correct future video. At an intermediate noise level \(t_k\), with \(\alpha_k,\sigma_k>0\),

\[x_k=\alpha_k f(c)+\sigma_k\epsilon,\qquad \epsilon\sim\mathcal N(0,I).\]

The exact conditional score is

\[s_k^*(x_k,c)=\frac{\alpha_k f(c)-x_k}{\sigma_k^2}.\]

Therefore, the Tweedie-denoised estimate is

\[\begin{aligned} \hat f_k^*(x_k,c) &=\frac{x_k+\sigma_k^2 s_k^*(x_k,c)}{\alpha_k}\\[6pt] &=\frac{x_k+\sigma_k^2\frac{\alpha_k f(c)-x_k}{\sigma_k^2}}{\alpha_k}\\[6pt] &=f(c). \end{aligned}\]

This holds for every \(x_k\). Even with different noisy states at different noise scales, the exact score always gives \(x_k\to f(c)\). Therefore,

\[\frac{\partial\hat f_k^*}{\partial x_k}=0.\]

This is the Jacobian of the clean-video estimate with respect to the current noisy state. Once conditioning \(c\) is fixed, changing \(x_k\) does not change the predicted future. The current state is irrelevant to the current clean estimate. More importantly, information carried through the denoising state from any earlier step cannot affect any later clean prediction. Every score evaluation, followed by the Tweedie formula, independently outputs the complete future \(f(c)\).

Information can propagate through the trajectory, but there is no path by which that propagated information can influence the computed solution \(f(c)\).

What about an approximate score?

\[s_k^\theta(x_k,c)=s_k^*(x_k,c)+e_k(x_k,c),\]

where \(e_k\) is the score error. The Tweedie-denoised estimate becomes

\[\begin{aligned} \hat f_k^\theta(x_k,c) &=\frac{x_k+\sigma_k^2 s_k^\theta(x_k,c)}{\alpha_k}\\[6pt] &=f(c)+\frac{\sigma_k^2}{\alpha_k}e_k(x_k,c)\\[6pt] &=f(c)+\lambda_k e_k(x_k,c), \end{aligned}\]

with \(\lambda_k=\sigma_k^2/\alpha_k\). Like the score, \(e_k\) is a vector field: it assigns a vector to each point in latent or video space. It depends on the noisy state \(x_k\) and contributes to the next denoising step. This provides a channel for recurrence.

As before, we can express this with the Jacobian:

\[\frac{\partial\hat f_k^\theta}{\partial x_k} =\lambda_k\frac{\partial e_k}{\partial x_k}.\]

If \(\partial e_k/\partial x_k\ne0\), information written into \(x_k\) by earlier steps can affect the computation at this step.

The Jacobian of a denoising step

An even more directly relevant object is the Jacobian of a single sampler step. For a numerical reverse-SDE update,

\[x_{k-1}=x_k+a_k x_k+b_k s_k^\theta(x_k,c)+\xi_k.\]

Holding the conditioning and sampled noise \(\xi_k\) fixed,

\[J_k:=\frac{\partial x_{k-1}}{\partial x_k} =(1+a_k)I+b_k\frac{\partial s_k^\theta}{\partial x_k}.\]

Since \(s_k^\theta=s_k^*+e_k\),

\[\begin{aligned} J_k&=(1+a_k)I+b_k\frac{\partial s_k^*}{\partial x_k} +b_k\frac{\partial e_k}{\partial x_k}\\[6pt] &=\left(1+a_k-\frac{b_k}{\sigma_k^2}\right)I +b_k\frac{\partial e_k}{\partial x_k}. \end{aligned}\]

The last term, \(b_k\,\partial e_k/\partial x_k\), is the part introduced by score error. Across several sampler steps, these state Jacobians multiply, just as in an RNN:

\[\begin{aligned} \frac{\partial x_0}{\partial x_k} &=\frac{\partial x_0}{\partial x_1} \frac{\partial x_1}{\partial x_2}\cdots \frac{\partial x_{k-1}}{\partial x_k}\\[6pt] &=J_1J_2\cdots J_k. \end{aligned}\]

In the perfect-score case, the Jacobians between denoising states can still be nonzero. For a later step \(j<k\),

\[\frac{\partial\hat f_j^*}{\partial x_k} =\underbrace{\frac{\partial\hat f_j^*}{\partial x_j}}_{=\,0} \frac{\partial x_j}{\partial x_k}=0.\]

The state-to-state Jacobian \(\partial x_j/\partial x_k\) can be nonzero, but the clean estimate’s Jacobian is zero. The result is zero regardless of how complicated the state-to-state dependence is.

This makes progress toward Amir’s question: “Why doesn’t composing score evaluations amount to serial computation?” With a perfect score, the state can carry information through the trajectory, but the map from the state to the task solution removes all dependence on that information.

How does SSH cover approximate scores?

This is where the low-score-error assumption enters. It would be useful to have a similar Jacobian argument for approximate scores, but low score error alone does not bound its derivative. Instead, SSH uses a convergence argument.

Let \(q_K(y\mid c)\) be the output distribution after \(K\) approximate-score sampling steps. In SSH’s formulation, the convergence bound has the form

\[\mathrm{TV}(q_K,p_{\mathrm{truth}}) \le C\left[ \frac{d_{\mathrm{int}}(\log K)^3}{K} +\epsilon_{\mathrm{score}}\sqrt{\log K} \right].\]

Here \(C\) is a theorem constant and \(d_{\mathrm{int}}\) denotes the discrete intrinsic dimension in the bound. The two terms account for:

  • Finite-step error: using finitely many denoising steps.
  • Score error: imperfect score estimation accumulated through the sampler.

When the intrinsic-dimension term stays bounded as input size grows, we can choose a fixed \(K^*\) large enough to make the first term small. Increasing \(K\) also increases the score-error term in this bound, so \(\epsilon_{\mathrm{score}}\) must be small enough for their sum to remain below the required threshold.

The formal premise is that one constant \(K^*\) satisfies this error threshold for every input size, together with the paper’s network, sampler, and moment assumptions. SSH then derives a constant-depth solution to the task. See Theorem F.2 and its interpretation.

02 / Something naive

Naive attempt
at guidance

CollapseExpand

Both experiments here run on the paper’s video models, so every number passes through a VAE and a ball tracker. Part 1 compares generating frames in order against generating them in one pass, and measures how much the one-pass sampler improves when it is handed the correct future. Part 2 tries to repair the one-pass model with guidance toward a simulated future.

Part 1 · Generation order

Autoregressive and bidirectional diffusion

Autoregressive diffusion gives lower errors than vanilla bidirectional diffusion across the tested reverse-step counts. Clean-latent oracle guidance reveals how much the bidirectional sampler can improve when given the correct future.

5 balls · 49 frames 896 holdout videos per seed Seeds 0–2

Both models use the same conditioning, targets, step counts, and metrics. We test 10, 20, 50, 100, and 200 reverse steps. Results are averaged over seeds, with no checkpoint averaging.

What we found

  1. Autoregressive diffusion outperforms vanilla bidirectional diffusion. This holds at every tested step count on rollout-5 error, exact-future error, and penalized rollout-5 error.
  2. Autoregressive accuracy does not improve steadily with more reverse steps. Its rollout-5 and exact-future errors rise slightly at 100–200 steps; penalized rollout-5 stays near 0.9 across the sweep.
  3. Clean-latent oracle guidance shows that bidirectional diffusion can use the correct target. It sharply lowers exact-future error and reaches the autoregressive model’s rollout-5 error around 50 steps. This guide uses the exact ground-truth future video. How the oracle is calculated →
  4. Penalizing invalid balls changes the oracle comparison. Oracle penalized rollout-5 is 13.07 at 10 steps and 12.50 at 20 steps. It beats autoregressive diffusion at 100 and 200 steps, reaching about 0.73 versus 0.90–0.93.

Rollout-5 and exact-future error

Download SVG
Autoregressive results by seed and comparison with vanilla bidirectional diffusion and clean-latent oracle guidance
Figure 1. The top row shows individual autoregressive seeds and their mean. The bottom row compares the three methods. Lower error is better.

Penalized rollout-5 error

Download SVG
Penalized rollout-5 error for autoregressive diffusion, vanilla bidirectional diffusion, and clean-latent oracle guidance
Figure 2. This metric penalizes missing, extra, disappearing, reappearing, and color-changing balls. Lower error is better.

Watch the same sample across methods

Generated video comparison

Choose a sample and a reverse-step count. Play, pause, and step through all five videos together. Every clip has 49 frames; playback follows the same frame position.

Frame 1 / 49

Loading the comparison…

Case 2 at 50 steps. Rollout-1 guidance produces missing, wrong-color, and extra balls. The stricter validity checks in “All metrics” score 19 of 220 ball/frame pairs.

Scores below each video use the repository’s original Rollout-5 and Rollout-1 metrics for the selected sample, seed 0. The plots above summarize the full holdout set across seeds 0–2.

Data behind the figures

Holdout summaries

Download the autoregressive results and the bidirectional evaluations used for penalized rollout-5.

Part 2 · Physics-based guidance

Rollout-1 and Rollout-5 guidance

Both guides improve a tuned Case 2 generation. In the earlier five-video tests, every tested variant worsens both mean errors. Below, we compare results that share the same vanilla controls and identify the settings that differed.

5 balls · 49 frames5 conditioning framesBidirectional DiT

Every result in Part 2 uses this video setting. Rollout-1 and Rollout-5 name the physics horizon used to build the guidance target. We evaluate the finished videos with both original repository metrics; lower is better. Ball-validity counts are separate.

Comparison 1 · Single-video tuning

Same generation and update cap; different strengths

Both guides improve both original metrics on Case 2. The tuned Rollout-1 result has lower errors than the saved cap-50 Rollout-5 result. All three clips share the single-video tuning experiment’s vanilla control.

Shared settings: seed 0, 50 denoising steps, guidance at step indices 25–49, and up to 50 latent updates per active step. Both guides apply direct gradient updates without checking and undoing proposals.

00000/828.mp4 · 50 denoising steps · seed 0 · single-video tuning

Vanilla bidirectional DiT

Open video ↗
Rollout-5 ↓1.4111Original repo metric
Rollout-1 ↓0.4316Original repo metric

Scorable Rollout-5 pairs: 55/220

Tuned rollout-1 guidance

Open video ↗
Rollout-5 ↓1.1459Original repo metric
Rollout-1 ↓0.3195Original repo metric

Scorable Rollout-5 pairs: 220/220

Rollout-5 guidance · cap-50

Open video ↗
Rollout-5 ↓1.3099Original repo metric
Rollout-1 ↓0.3817Original repo metric

Scorable Rollout-5 pairs: 204/220

Frame 1 / 49

Loading the three videos…

SettingRollout-1 guidanceRollout-5 guidance
Guidance strength0.70 → 0.20; final step set to 0.100.20 → 0.06
How this output was selectedLowest original errors among 21 final-batch variantsSelected for visual quality during the sweep; original Rollout-5 later measured 7.2% below vanilla
Tests on more videos with these settingsNoneNone

The strengths were tuned separately. This comparison shows what the two saved configurations achieved on one generation; changing the guidance horizon was not the only difference. Only two final videos survive from the 60-run Rollout-5 sweep, so the original metric cannot rank all 60 settings.

Comparison 2 · Earlier five-video tests

Same five controls and guidance budget; different update rules

The Rollout-1 pilot and all six late Rollout-5 variants reuse the same five vanilla videos, verified by file hashes. They share seed 0, 50 denoising steps, guidance at step indices 40–49, and up to 12 accepted updates per step. Each proposal tries strengths 0.1, 0.03, 0.01, and 0.003.

Every tested variant raises both mean errors. Vanilla scores 0.4253 on Rollout-1 and 1.4163 on Rollout-5. The Rollout-1 pilot scores 0.4332 / 1.5004; the Rollout-5 variants score 0.4522–0.4922 / 1.4540–1.5192.

Neither pilot improves the mean original metrics

Download SVG ↓
Original Rollout-1 and Rollout-5 means on five shared controls: the Rollout-1 pilot and all six Rollout-5 variants worsen both metrics
Figure 3. R1 and R5 identify the guidance horizon. All clips use the same original evaluator and exclude the five conditioning frames. Rollout-1 scores for the 30 saved Rollout-5-guided clips were calculated for this comparison; their recalculated Rollout-5 scores match the report. Per-video data and hashes · Means.

What differed between the pilots?

ChoiceRollout-1 pilotRollout-5 variants
Frames used for each updateChoose the frame to update with the largest current errorCombine losses from all five-frame windows for which ball positions and velocities can be estimated
TargetEncode a reference with one frame replaced by a one-step physics predictionEncode a five-step physics reference, or use the latent change caused by that reference
Where the loss is measuredLatent slice for the chosen frameWhole five-frame window or its endpoint
Ball detectionEarlier detectorSeveral detector versions; strict detection rejects windows starting from a frame with an extra ball

These pilots accept an update only when their guidance score improves. The Case 2 runs above instead use direct updates and select the smallest error above a threshold. The two experiment stages therefore test different procedures.

In the plot, “direct target” means the encoded physics reference. “Delta” subtracts the change caused by encoding and decoding the current video. The separate ball-validity counts were produced by different tracker versions, so we compare the common original metrics here. Download the shared settings and each run’s configuration.

Additional evidence · Same video setting

Tests without a matching guidance-horizon comparison

The following experiments also use five balls and 49 frames. They test update budgets, more videos, or alternative targets, but do not provide a matched Rollout-1 versus Rollout-5 comparison.

Rollout-1 only: larger update budgets did not help

Download SVG ↓
Case 2 update-cap sweep: caps 75 through 500 worsen both original rollout errors compared with cap 50; cap 51 improves errors slightly but loses complete ball validity
Figure 4. Single-video tuning, Case 2 at 50 denoising steps. The sweep uses strength 0.70 → 0.20, including 0.20 on the final step; the tuned video above uses 0.10 on that step. Caps 75, 100, 200, and 500 all worsen both errors relative to cap 50. Cap 51 lowers both errors slightly, with 214/220 scorable transitions. Cap 500 uses ten times the guidance updates of cap 50. Download data.

Rollout-1 only: the ten-video test worsened both metrics

Download SVG ↓
Ten-video means at 50, 100, and 200 denoising steps: the 75-update guidance variant increases original Rollout-1 and Rollout-5 error and reduces scorable transitions
Figure 5. No matching Rollout-5 run was conducted. This is the 75-update configuration shown in the main video viewer, with strength 0.70 → 0.20 and guidance starting at step index 25. Both original metrics worsen on all ten videos at each active step count. At 50 steps, mean Rollout-1 rises from 0.2317 to 0.6175, and Rollout-5 from 0.9368 to 1.8875. Its Case 2 uses different starting noise and a different vanilla control from the tuned example above. The 10- and 20-step runs apply zero guidance updates and are omitted here. Download data.

Rollout-5 only: the larger holdout test

For five balls and 49 frames, the initial physics-derived targets barely change original Rollout-5. The exact-future oracle lowers it by 34.92%.

Latent correctionOriginal Rollout-5 improvement
Random correction of matched size+0.16%
Encoded five-step physics target+0.08%
Latent change caused by the physics target+0.12%
Exact-future oracle+34.92%

896 videos per seed, seeds 0–2; percent improvements averaged over 50, 100, and 200 denoising steps. Positive values mean lower error. These initial physics targets use the evaluation’s exact initial state; the oracle uses the correct future. No corresponding Rollout-1 target was tested in this holdout experiment. Download this video setting’s data.

What would isolate the effect of guidance horizon?

A matched test would change only the physics horizon, keeping the videos, starting noise, update rule, strength schedule, and update budget fixed. The saved experiments provide shared-control comparisons, but do not make that single change.

Reports and data

The archives document all tested settings. The full Rollout-5 report also includes other ball counts and video lengths; the comparison above is restricted to five balls and 49 frames. All videos shown here are saved experiment outputs.

03 / State diffusion

A cleaner set-up,
(position-velocity) state diffusion

CollapseExpand

We retrained the billiards experiment from scratch on positions and velocities instead of video. The gap between one-pass and block-by-block generation reproduces at 1/30 of the parameters, with no VAE and no tracker anywhere in the loop.

1, 2, 3, 5 balls · 25–49 frames 1,024 held-out clips Seeds 0–2 6.85M parameters

The seriality-gap paper’s released checkpoints generate pixels, so every number they produce passes through a VAE encoder, a VAE decoder, and a ball tracker. Those stages carry their own error, and it is not separable from the model’s: the repository’s own check puts the local-error floor of that pipeline at 0.19 world units, which is close to the 1-ball values the paper reports. These models predict the state directly and the metric is exact to about 1e-6, so nothing is hidden under a floor. The datasets were also rebuilt: every clip comes from its own independent simulation, where the paper’s generator takes up to 1,000 overlapping clips from each of about twenty simulations, making dataset diversity a hidden variable along the clip-length axis the experiment measures. “The paper” below means the seriality-gap paper, not SSH. Source: experiments/statebench/RETRAIN_REPORT.md.

3.1 · The physics simulator

How the ground-truth simulator works

Every clip in this section, and every error number, comes from one simulator: experiments/statebench/statebench/sim/solver.py, the seriality-gap paper’s own. This is what it does.

The world

A square box, 10 by 10 units. Inside it, 1, 2, 3 or 5 balls, all of radius 0.7 and equal mass. No gravity, no friction, no spin: between contacts a ball moves in a straight line at constant speed. Contacts are perfectly elastic. Speeds in the data have a median of 8 units per second (5% to 95%: 2.3 to 14.7), so in one frame at 15 frames per second a typical ball moves 0.55 units, about 0.8 of its radius. What the data records, once per frame and per ball, is four numbers: the centre position (x, y) and the velocity (vx, vy) at that instant. A 49-frame, 5-ball clip is therefore 49 × 5 × 4 = 980 numbers, and that is exactly what the state models take in and put out.

How time advances

The simulator has no time step. It is event-driven: it jumps from one contact to the next and moves every ball in a straight line in between. For one frame interval (1/15 s) it does the following.

  1. From the current positions and velocities, compute when each possible contact would happen: for every pair of balls, the time at which their centres are 2 × 0.7 = 1.4 units apart (the root of a quadratic in time); for every ball and each of the four walls, the time at which its centre is 0.7 from the wall (a linear equation).
  2. Take the earliest of these times. If it is after the end of the frame interval, move every ball in a straight line to the end of the interval, record the frame, and stop.
  3. Otherwise move every ball in a straight line to that time, change the velocities of the balls involved by the rule for that contact (next paragraph), and go back to step 1 from the new state.

Because the contact times are solved for exactly and the motion between them is exact, the recorded frames carry no integration error. Any number of contacts can happen inside one frame interval, each at its own time; the frame only shows where the balls ended up.

The two contact rules

Download PNG
Left: a ball hitting a wall, its velocity component into the wall reversed and the component along the wall kept. Right: two equal balls in contact, exchanging their velocity components along the line joining their centres.
Figure 6. Left: at a wall, the velocity component perpendicular to the wall changes sign and the component along the wall is unchanged. Right: when two balls touch, they exchange the components of their velocities along the line joining their centres (the same as reflecting their relative velocity about that line) and keep the components perpendicular to it. Both rules leave the total momentum and the total kinetic energy of the balls unchanged, apart from the momentum a wall absorbs.

What happens between two recorded frames

Download PNG
Left: the five balls of a held-out clip at frames 6 (dashed) and 7 (filled). Right: zoom on the lower-left corner, showing the cyan ball's exact path during the frame interval: it bounces off the left wall 0.0258 s after frame 6 and collides with the green ball 0.0608 s after it, so its path is two straight segments joined at two corners while the two recorded frames alone suggest one straight move.
Figure 7. Held-out clip 0, frames 6 and 7. In the 1/15 s between them the cyan ball bounces off the left wall (0.0258 s after frame 6) and then collides with the green ball (0.0608 s after frame 6). The two recorded frames show only the end points; the straight line between them (black dashed) is not the path the ball took, and the recorded velocity at frame 7 is the velocity after both events. In 5-ball, 49-frame clips, 82% of the ball-frame intervals contain no contact, 15% contain one, and 3% contain two or more; of the intervals with a contact, 17% have more than one. A clip has on average 14.4 ball-ball collisions and 20.7 wall bounces in its 44 generated frame intervals, and 99.6% of clips contain at least one interval with two contacts for the same ball.

Four frame intervals in slow motion

Open video
Figure 8. Held-out clip 0 from frame 5 to frame 9, the exact motion sampled 30 times per frame interval and played at 1/30 of real speed. Dashed outlines mark the recorded frames as they are reached; the label appears at the moment of each wall bounce or collision. Frame 6 to 7 is the interval of Figure 7.

How the simulator scores a generated clip

The same simulator produces the error numbers. The global error Δx(GT) is the mean distance, over balls and over frames 5 to 48, between a generated ball centre and the true one. The local error Δx(k) asks whether the generated clip obeys the physics from one moment to the next: take the generated frame t−k, estimate each ball’s velocity from its position change since frame t−k−1 (the paper’s estimator; the model’s own velocity numbers are not used, so the score depends on positions only), run the simulator k frames forward from that state, and measure the distance to the generated frame t. The mean over balls and over t = 5 … 48 is Δx(k); the paper uses k = 1 and k = 5. A clip that follows the physics exactly scores 0 on Δx(1) and about 0.02 on Δx(5) (the small floor comes from estimating velocities from positions when a contact falls inside the interval). The true clips score exactly that.

The five-step local error on one generated clip

Download PNG
A clip generated by the bidirectional model, frames 13 to 18, with the simulator's five-frame prediction for the green ball drawn as black squares and a red arrow from the predicted position to the generated one, labelled 1.58 units.
Figure 9. Frames 13 to 18 of a clip generated by the bidirectional model. Coloured: the model’s positions (dashed outline at frame 13, solid at frame 18). Black squares: the simulator run five frames from the model’s own frame 13. The red arrow is this ball’s contribution to Δx(5) at frame 18, 1.58 units: the model moved the green ball down while, from its own frame-13 state, the physics says it should have kept going up. Δx(5) is the mean of such distances over all balls and frames.

One property of this score matters later: every distance in a clip scales with the balls’ speed, so a generated clip whose balls move 30% slower than the data scores about 30% lower on Δx(5) without following the physics any better. A fair reading of Δx(5) needs the sample’s mean speed next to it.

What we found

  1. Generating in order is four times more physically accurate. At 5 balls and 49 frames the block-4 model reaches 0.223 local error against the bidirectional model’s 0.885. The paper’s video models split 0.607 against 0.867 at about 200M parameters; these models have 6.85M.
  2. Finer generation order keeps helping, all the way down. 0.885 for all frames at once, 0.223 at 4 frames per block, 0.110 one frame at a time. No bidirectional model of any depth, width, step count, or noise schedule we tested comes close.
  3. What hurts one-pass generation is dependent collisions, not clip length. With one ball the future is a closed-form function of the given frames, and the two architectures become nearly equal (0.0300 against 0.0229) and flat in clip length. With 5 balls the gap grows from 2.94× at 25 frames to 3.98× at 49.
  4. Score error per predicted number does not grow with clip length. At every ball count it is flat or slightly falling. The total score error grows as the square root of the output size, which is exactly what a fixed per-number error produces on a larger output — not a length-dependent breakdown in score accuracy.

What the model works on

Figure 10. Four held-out clips per ball count. Left: truth. Middle: the bidirectional model. Right: the block-4 autoregressive model — this panel is empty at 2 and 3 balls, which have no block-4 run. Each model receives frames 0–4, marked “given”, and generates frames 5–48 at 50 denoising steps per block. Playback is 15 frames per second. Every clip is rendered from the saved generations behind the numbers below: seed 0, the 25,000-step checkpoint, sampling seed 0.

Balls of radius 0.7 collide elastically in a 10 × 10 box at 15 frames per second. The model sees only numbers: 49 frames × 5 balls × (x, y, vx, vy) = 980 values per clip. No pixels, no VAE, no tracker. It is a small transformer with one token per (frame, ball), 8 layers, width 256, 6.85M parameters, trained with flow matching.

The datasets

Each dataset is built by the paper’s own generator: 400 frames simulated per seed, the first 200 discarded as burn-in, at most one accepted window kept per seed. Initial conditions, solver, window filters, collision-density limits, and the held-out collision-count quotas are the original code, unchanged.

Set, every cellClipsIndependent simulations
Training20,00020,000
Held-out1,0241,024

The burn-in keeps the state distribution equal to the original data’s. First-frame ball speeds, 5 balls at 49 frames:

Quantile5%25%50%75%95%
This data2.285.328.0410.9014.66
Original data2.275.308.0110.8814.66

Without the burn-in every clip would start from a freshly sampled initial condition, where 100% of ball speeds lie in [8, 10] against 18% in equilibrated data. Held-out sets keep the original repository’s balancing across ball-ball collision counts; one ball has no ball-ball collisions and no balancing.

Main result at 5 balls

Download PNG
Three bar panels — global error, one-step local error, five-step local error — comparing the paper's video models with the bidirectional and block-4 state models
Figure 11. All 1,024 held-out clips, 49 frames, the paper’s metric definitions, mean over 3 seeds. Grey bars are the paper’s published video-model values, at about 200M parameters against these models’ 6.85M. Lower is better.

Global error Δx(GT) is the mean distance between generated and true ball centres over frames 5 to 48. Local error Δx(k) runs the exact simulator k frames forward from the generated frame t−k and compares it with the generated frame t, with velocities read off the generated positions by the paper’s finite-difference estimator.

The bidirectional model reaches 0.885 on the five-step local error. The block-4 model — the paper’s autoregressive scheme applied to states, 4 frames generated per block, earlier frames held fixed — reaches 0.223, four times more physically accurate, and also has the lower global error.

More denoising steps stop helping at about 20

Download PNG
Two line panels against number of denoising steps: local error falls until about 20 steps and is then flat; the bidirectional model's global error is lowest at one step and rises
Figure 12. 5 balls, 49 frames, mean over 3 seeds. The lighter lines use the model’s own generated velocities instead of reading them off positions; comparisons with the paper must use the paper’s estimator, the dark lines.

Left: the local error falls until about 20 steps and is flat from 20 to 200, the paper’s Observation 4. No number of denoising steps brings the bidirectional model near the block-4 model.

Right: the bidirectional model’s global error is lowest at 1 step and rises with more steps. One step returns roughly the conditional mean of all futures consistent with the given frames — close to the truth on average, but not a physically consistent trajectory. The global error therefore does not measure physical accuracy; the local error does.

Where the error comes from

Download PNG
Distance between generated and true ball positions plotted against frame index, for one and fifty denoising steps, with a straight-line no-collision reference
Figure 13. Mean distance between generated and true positions at each frame, over the 1,024 held-out clips. The grey line shows how fast error grows with no collision handling at all: straight-line motion from frame 4.

The bidirectional model is accurate at the first generated frame and drifts to the level of unrelated positions in the box by about frame 30.

The 1-step curve lies below the 50-step curve, which looks backwards but is expected. One denoising step returns a blend of trajectories rather than any one of them; being the mean, it minimises squared distance to the truth while being physically wrong. Fifty steps return one physically consistent sample, and because the dynamics are chaotic a single valid future ends up as far from the true one as any other valid future.

Held-out local error against training step for the bidirectional and block-4 models, all three seeds drawn, both falling monotonically to 25,000 steps
Figure 14. Held-out local error on 128 validation clips, measured every 5,000 steps, all three seeds drawn. Both panels fall monotonically to the 25,000-step checkpoint, and the three seeds of each configuration lie on top of one another. Every result in this section uses the checkpoint at exactly 25,000 updates, never the best validation score.

Control: one ball

Download PNG
Bar comparison of five-ball and one-ball local error for the bidirectional and block-4 models: at one ball both are far lower and nearly equal
Figure 15. 49 frames, mean over 3 seeds. With one ball there are no ball-ball collisions, only wall bounces, and the future is a closed-form function of the given frames.

The paper’s control removes the serial structure. Both models become far more accurate, and — the substantive point — the two architectures become nearly equal: 0.0300 bidirectional against 0.0229 block-4 at 49 frames, a ratio of 1.31. Global error agrees: 0.0434 against 0.0399.

That matches the paper’s Observation 2, which reports 0.246 and 0.220 for its two video models at one ball and 49 frames, a ratio of 1.12. The gap that exists at 5 balls does not exist at 1 ball.

Horizon sweep

Download PNG
Local error against clip length: at five balls the bidirectional curve rises and the block-4 curve is flat; at one ball both are flat and nearly equal
Figure 16. Local error against clip length, 5 balls on the left and the 1-ball control on the right. Filled circles are the state models, mean over 3 seeds with the standard deviation across seeds as error bars; dotted squares are the paper’s published video-model values.

5 balls: the bidirectional model degrades with clip length, 0.684 at 25 frames to 0.885 at 49, tracking the paper’s video curve (0.670 to 0.883). The block-4 model is flat at 0.223 to 0.239, as the paper’s autoregressive video model is flat at about 0.61. Observations 1 and 3 reproduce, at 1/30 of the parameters.

1 ball: both models are flat, and so is the gap between them. The bidirectional model goes from 0.0371 at 25 frames to 0.0300 at 49 — slightly down, not up. Clip length by itself is not what hurts one-pass generation; dependent collisions are.

All ball counts

Download PNG
Local five-step error and global error against clip length, one curve per ball count, for the bidirectional models
Figure 17. Bidirectional models, local five-step error and global error, mean over 3 seeds. Error rises with ball count at every length, and with length at every ball count above one.

Local five-step error, bidirectional models, mean over 3 seeds with the standard deviation across seeds:

Balls25 frames33 frames41 frames49 frames25 → 49
10.0371 ±0.00170.0343 ±0.00040.0316 ±0.00030.0300 ±0.00010.81×
20.1774 ±0.00300.2653 ±0.00430.2970 ±0.01200.3433 ±0.00811.94×
30.3833 ±0.01840.5474 ±0.01600.6164 ±0.01450.6432 ±0.00861.68×
50.6841 ±0.00950.8396 ±0.02910.8846 ±0.01560.8851 ±0.01561.29×

Block-4 models, with the bidirectional-to-block-4 ratio in parentheses. The gap grows with clip length at 5 balls and is flat at 1 ball:

Balls25 frames33 frames41 frames49 frames
10.0320 ±0.0007 (1.16×)0.0272 ±0.0004 (1.26×)0.0258 ±0.0013 (1.23×)0.0229 ±0.0005 (1.31×)
50.2326 ±0.0044 (2.94×)0.2388 ±0.0028 (3.52×)0.2332 ±0.0049 (3.79×)0.2226 ±0.0026 (3.98×)

Conditional score error

Total score RMSE from the score-error audit run on these checkpoints: 25,000-step weights, 1,024 held-out clips, 32 uniform noise levels, 8 noise draws, float32 with TF32 disabled. “Generated frames” is clip length minus the 5 conditioning frames.

Balls20 generated frames283644
119.13 ±0.6023.15 ±0.2225.91 ±0.5327.91 ±0.47
252.50 ±0.3462.81 ±0.5671.35 ±0.7777.59 ±0.07
365.56 ±0.5980.99 ±0.6891.72 ±0.3398.15 ±0.43
587.68 ±1.24103.28 ±2.02112.24 ±2.10120.39 ±1.17

The same audit, as mean squared score error per predicted number:

Balls20 generated frames283644
14.5754.7834.6644.426
217.22917.61617.68017.103
317.91219.52219.47418.247
519.22319.05117.50116.471

Per-number score error is flat or slightly declining with length at every ball count, one ball included. The total grows only because there are more numbers to predict. From 20 to 44 generated frames the output grows by a factor of 2.2, so a flat per-number error predicts √2.2 = 1.48× growth in the total; measured growth is 1.46× at 1 ball, 1.50× at 3 balls, and 1.37× at 5 balls.

This matters for the audit’s question. A total score error that grows as the square root of the output size, with no increase per number, is what a model whose per-number score error is fixed and positive produces on a larger output. It is not evidence of a length-dependent breakdown in score accuracy.

Sampling accuracy

Denoising-step budgets of 1, 2, 4, 8, 16, 32, 64, and 128, shared initial-noise seed, all 1,024 held-out clips. Best mean local error over all tested budgets:

Balls20 generated frames283644
10.03380.02970.02730.0256
20.17550.25240.29720.3351
30.38950.53980.62520.6288
50.68370.83110.89260.8927

A successful global prediction has mean position error at most 0.1 world units; a successful local prediction has five-frame physics error at most 0.05. The 1-ball models meet both thresholds at a fixed budget that does not grow with length: 2 steps for the global criterion at every length, and 32 steps for the stricter local criterion at 28, 36, and 44 generated frames. The 20-frame case needs 128 steps for the local criterion, the one place a shorter clip is harder — its best mean local error, 0.0338, is also the worst of the four. No multi-ball cell reaches 90% on either criterion at any tested budget, and adding steps does not fix it.

Block size: finer generation order, better physics

Download PNG
Local error and global error against the number of frames denoised together, falling monotonically from 44 frames at once down to one frame at a time
Figure 18. Block size is how many frames are denoised together before being frozen so the next block starts. 44 means all generated frames at once, which is the bidirectional model; 1 means one frame at a time. 5 balls, 49 frames, mean over 3 seeds.
Frames per blockLocal error Δx(5)Global error Δx(GT)
44 (bidirectional)0.8851 ±0.01562.928
220.7459 ±0.01582.793
110.4730 ±0.02142.600
40.2226 ±0.00262.358
10.1100 ±0.00442.115

Both metrics fall monotonically as blocks get finer — the paper’s Observation 3, with a much wider spread than in pixel space, where its models span 0.88 down to 0.61. Block-causal models are cheaper at inference than the bidirectional one only when blocks are large: block 1 costs 44 times as many model calls at the same number of denoising steps.

Block 1 also shows the step saturation: 0.401 at 1 denoising step, 0.133 at 5, 0.116 at 20, 0.115 at 50.

Depth buys more than width

Download PNG
Local error against depth at fixed width and against width at fixed depth: depth keeps improving to 32 layers while width saturates immediately
Figure 19. Bidirectional models, 5 balls, 49 frames, mean over 3 seeds. Depth 64 and width 1024 were not rerun.
Depth at width 256Local errorWidth at depth 8Local error
41.1566 ±0.01711281.0824 ±0.0233
80.8851 ±0.01562560.8851 ±0.0156
160.7646 ±0.00555120.8492 ±0.0117
320.7171 ±0.0109——

Depth gains keep coming to 32 layers but flatten (0.885 to 0.765 to 0.717), while width saturates almost immediately (0.885 to 0.849 for a 4× parameter increase). Depth buys more than width per parameter, the paper’s Observation 5.

No bidirectional model of any size tested approaches the block-4 model’s 0.223, let alone block 1’s 0.110. More layers do not substitute for generating in order.

The noise schedule does not change the conclusions

Download PNG
Bars comparing shift 1 and shift 5 noise schedules for the bidirectional and block-4 models: shift 5 is worse for both and the ordering is unchanged
Figure 20. 5 balls, 49 frames, 3 seeds each. The paper’s video models use Wan’s noise schedule (shift 5), which concentrates training and sampling effort at high noise; this bench uses uniform noise levels (shift 1).
ModelShift 1Shift 5
Bidirectional0.8851 ±0.01561.0442 ±0.0115
Block 40.2226 ±0.00260.2828 ±0.0023

Shift 5 makes both models worse and leaves the ordering and the size of the gap unchanged (3.98× against 3.69×). The schedule choice is not what produces the results above.

Horizontal bars of training minutes per run by ball count and architecture, from 7 minutes for one-ball bidirectional to 50 minutes for five-ball block-4
Figure 21. Training time for one 49-frame run, 25,000 steps, on one A100. Data generation is CPU-only and takes seconds to minutes per cell, with the one exception below.

7 minutes for 1-ball bidirectional, 22 minutes for 5-ball bidirectional, 10 and 50 minutes for the corresponding block-4 models. Block-causal training is slower because the paper’s scheme doubles the sequence and needs an explicit attention mask, which prevents the fused attention kernel; at one ball the sequence is short enough that this costs almost nothing.

Side note: predicting how long a balanced held-out set takes to build

The original repository balances each held-out set across ball-ball collision counts. Under one clip per simulation, the rarest bin must be found by rejection sampling rather than harvested from a single collision-heavy simulation, and the cost varies enormously by cell.

The collision-count distribution has a geometric tail, so a short sample predicts the rare bins. 24,000 fresh simulations for the 2-ball 49-frame cell, 82 seconds on 8 cores:

Ball-ball collisions45678
Accepted clips6,7191,094129204

Consecutive ratios are near-constant at 0.140. Fitting on the well-populated bins and extrapolating:

CollisionsPredicted rateSimulations for a 171-clip bin
77.5e-4227,000
81.1e-41.6M
91.5e-511.5M

Checked against the live build: the model predicts about 119 clips in the 9-collision bin after 8.0M simulations; the build had about 100. Roughly 20% accurate, from 82 seconds of sampling.

A single-rate Poisson process would have a factorial tail, falling off much faster than this. A constant ratio is what a Poisson mixture gives: mixing over a Gamma-distributed rate yields a negative binomial, whose tail is geometric. Physically that fits, since how often two balls meet depends on their configuration, so the collision rate varies between simulations. This is an inference from the shape, not verified against the simulator.

CellSimulationsWall time
Any 1-ball cell~24,00013 s
5-ball, 49 frames23,0004 min
3-ball, 49 frames314,00021 min
2-ball, 41 frames1.0M33 min
2-ball, 49 frames~11M~2.5 h

Practical rule: before building a balanced cell, run about 25,000 simulations, fit the ratio on the bins with good counts, and extrapolate. A 6,000-simulation smoke test whose rarest bin has a single observation gives an unusable estimate.

Scope and data

What was run

102 training runs: bidirectional at 1, 2, 3, and 5 balls and block-4 at 1 and 5 balls, each at 25, 33, 41, and 49 frames with three seeds. 48 score audits, 48 sampling sweeps, and 384 physics evaluations on the bidirectional checkpoints. All 16 datasets rebuilt and validated. A further 30 runs cover the block-size, depth, width, and noise-schedule sweeps at 5 balls and 49 frames. Not rerun: depth 64 and width 1024, the two most expensive families, and the memorisation gates.

Per-seed numbers

One row per training run: global error, one-step local error, and five-step local error at 50 denoising steps, read from each run’s saved evaluation. These are the numbers behind the horizon, ball-count, block-size, depth, width, and schedule results above. The score-error and sampling-accuracy tables come from the separate score-error audit.

03.5 / Score-error audit · side path

Actually calculating
GT score error

CollapseExpand

A side path off the main sequence: instead of measuring generated trajectories, we measure the models’ score itself against the exact score, which the simulator makes available. Summed over a trajectory the error grows with clip length. Per state component it is flat. The theorem’s low-score-error assumption is exactly the per-component error shrinking like 1/√d, and that shrinkage is absent at every ball count. Averaged over the noise levels a run of \(K\) denoising steps evaluates the network at, rather than over a fixed grid, it grows with \(K\).

1, 2, 3, 5 balls · 20–44 generated frames 1,024 held-out clips Seeds 0–2 32 noise levels × 8 noise draws 48 audits · 384 physics evaluations

This audits the same state models as section 03, at their 25,000-update checkpoints. A state component is one ball’s x position, y position, horizontal velocity, or vertical velocity in one frame; a clip of B balls and F future frames has \(d=4BF\) of them. Network evaluations run in float32 with TF32 disabled; squared residuals and their sums accumulate in float64. The noise levels 0.001, 0.003 and 0.01 are separate diagnostics and are excluded from every primary average. Source: experiments/score-error-audit/REPORT.md.

What we found

  1. The total error grows; the error per state component does not. From 20 to 44 generated frames the output grows by 2.2×. The error summed over a trajectory grows by 1.37–1.50×, and √2.2 = 1.48. The error per state component changes by 0.93–1.01×.
  2. The assumption is equivalent to a trend we can measure, and the trend is absent. The two quantities are related by an exact identity, \(\varepsilon_{\mathsf{score}}(d)=\sqrt{d}\,a(d)\), so “\(\varepsilon_{\mathsf{score}}\le C\) for every \(d\)” is the same claim as “\(a(d)\le C/\sqrt d\)”. Keeping the total fixed across the tested range needs the per-component error to fall to 0.67× of its starting value. It changes by 0.93–1.01×.
  3. The measured growth is the pure √d rate. Fitting \(\log\varepsilon_{\mathsf{score}}=\text{const}+p\log F\) gives \(p\) = 0.482, 0.499, 0.518 and 0.399 at 1, 2, 3 and 5 balls, with a largest residual of 0.021 in log space. \(p=0.5\) is exactly √d growth; \(p=0\) would mean the total had stopped growing.
  4. The paper’s one-ball control reproduces. At one ball, from 20 to 44 generated frames, the five-frame physics error goes 0.036 → 0.028 and the score error per predicted number 4.575 → 4.426. Both hold up; only the total grows, and it grows because there are more numbers to predict.
  5. Score accuracy and generation quality coincide on every cell of the grid. One ball is the only case with a small per-component score error, and the only case where quality holds up with length. So these measurements cannot separate a depth limit from the model failing to learn the score — both predict the same degradation.
  6. The score error the bound uses rises with the denoising step count, while the clips get more accurate. Averaged over the noise levels a run of \(K\) denoising steps actually evaluates, rather than over the audit’s fixed 32-level grid, each four-fold increase of \(K\) multiplies the score error by about 3, at every ball count. Over the same range the five-frame local position error falls and then flattens.

How the score error is calculated

The exact score is available because the conditioning frames determine one future and the noise we add has a known distribution. The simulator supplies that future; differentiating the Gaussian log density gives the score. The simulator is never differentiated.

1. Take the true future from the held-out simulation

Each held-out clip stores the simulator’s positions and velocities for every ball at every frame. The first five frames are the conditioning, \(c\). With the simulation parameters fixed, they determine one future, \(y^*(c)\), holding \(d=4BF\) state components. Positions are centred at 5 and divided by 2.5, velocities divided by 6.25; every score below is with respect to these standardized variables.

2. Add noise and write down its exact conditional score

We draw independent standard Gaussian noise for those \(d\) numbers and form a noisy future, leaving the five conditioning frames clean:

\[\epsilon\sim\mathcal N(0,I_d),\qquad z_\sigma=(1-\sigma)y^*(c)+\sigma\epsilon,\qquad 0<\sigma<1.\]

Because \(y^*(c)\) is fixed once \(c\) is given, the noisy future is Gaussian with a known mean and variance, \(p_\sigma(z\mid c)=\mathcal N\!\left(z;(1-\sigma)y^*(c),\sigma^2 I_d\right)\). Differentiating its log density with respect to \(z\) gives the score, and at the sample we just built it collapses to the noise we drew:

\[s^*(z\mid c,\sigma)=\frac{(1-\sigma)y^*(c)-z}{\sigma^2},\qquad s^*(z_\sigma\mid c,\sigma)=-\frac{\epsilon}{\sigma}.\]

This exactness depends on one future per conditioning. If the same conditioning allowed several futures, the noisy distribution would be a mixture and one example’s noise would not give its score.

3. Convert the model’s output to a score and subtract

These networks were trained with flow matching: they predict \(\hat v_\theta(z_\sigma,c,\sigma)\) against the target \(v^*=\epsilon-y^*(c)\). Under this noise convention the score implied by a predicted flow, and its error, are

\[\hat s_\theta=-\frac{z_\sigma+(1-\sigma)\hat v_\theta}{\sigma},\qquad \hat s_\theta-s^*=-\frac{1-\sigma}{\sigma}\left(\hat v_\theta-v^*\right).\]

So the audit computes the squared score error straight from the flow residual, which is algebraically the same as building both score vectors and subtracting. Only the \(F\) future frames enter the sum. Each noisy input costs one network evaluation; the separate sampling experiment runs the full denoising loop instead.

4. Average, and report two numbers

For each trained model we evaluate 1,024 held-out clips at the 32 levels \(\sigma_k=(k+1/2)/32\), with eight independent noise draws per clip reused across levels. Score error for the full trajectory sums the squared errors over all \(d\) state components, averages over clips, draws and levels, then takes the square root. Score error per state component averages over the \(d\) components as well before the square root, so it is the first divided by \(\sqrt d\):

\[E_r=\sqrt{\operatorname{mean}\left[\sum_{j=1}^{d}\left(\hat s_{r,j}-s_j^*\right)^2\right]},\qquad R_r=\frac{E_r}{\sqrt{d}}.\]

Both take a square root; their only difference is whether squared errors are summed or averaged over components. Tables and frame-count curves average \(E_r\) or \(R_r\) across the three training seeds, with error bars showing the standard deviation across seeds. Noise-level plots use one level per point. The full-trajectory number uses the same unnormalized L2 norm as the assumption it tests.

The total grows, the per-component error does not

Download PNG
Two panels against generated frame count: score error summed over a trajectory rises at every ball count, while score error per state component stays flat
Figure 22. Left: squared score errors summed over all state components in the future trajectory. Right: averaged over those components. Both average over held-out clips, noise draws and the 32 primary noise levels before taking the square root. Lines average 3 training seeds; shading and error bars show ±1 SD across seeds. Both vertical axes are logarithmic.

From 20 to 44 generated frames, the total changes by 1.46× at one ball, 1.48× at two, 1.50× at three and 1.37× at five. The per-component error changes by 0.98×, 1.00×, 1.01× and 0.93× over the same range.

BallsGenerated framesSeedsFull trajectory (RMS L2)Per state component (RMSE)
120319.132.138
128323.152.187
136325.912.159
144327.912.104
220352.504.151
228362.814.197
236371.354.205
244377.594.136
320365.564.232
328380.994.418
336391.724.413
344398.154.272
520387.684.384
5283103.284.364
5363112.244.183
5443120.394.058

Means over 3 training seeds. Download the score summary, which also carries the per-seed values and the per-component figure as score_rmse_per_component.

The same totals, level by level

Download PNG
Score error for the full trajectory against noise level, one panel per ball count
Figure 23. At each noise level, squared score errors are summed over all state components, averaged over held-out clips and noise draws, then square-rooted. Lines average training seeds; bands show ±1 SD across seeds. The table above additionally averages over noise levels before the square root.

Per state component, level by level

Download PNG
Score RMSE per state component against noise level, one panel per ball count
Figure 24. Squared errors averaged over state components as well as clips and draws, then square-rooted: each model’s value is its full-trajectory error divided by \(\sqrt{4BF}\). Lines average training seeds; shading shows ±1 SD. The assumption being tested uses the unnormalized norm of the previous figure.

How much clips differ from one another

Download PNG
Median held-out clip error against generated frame count, with bands spanning the 16th to 84th percentiles
Figure 25. Per clip and model, squared errors are summed over state components and averaged over the eight noise draws and the noise levels, then square-rooted; the right panel also divides by \(\sqrt{d}\). Errors are averaged across training seeds for each clip. Lines show the median clip; bands span the 16th to 84th percentiles, the middle 68% of the 1,024 clips. These are spread bands, not confidence intervals for the mean.

Clip-to-clip spread by noise level

Download PNG
Median held-out clip error against noise level with 16th-to-84th-percentile bands, one panel per ball count
Figure 26. The same per-clip construction at each noise level. These median curves describe a typical clip and differ from the dataset-wide RMS curves above. Download the spread statistics: mean, SD, median and percentile bounds.

No step budget up to 128 fixes the multi-ball case

Download PNG
Five-frame local physics error against denoising step budget, one panel per ball count
Figure 27. Budgets of 1, 2, 4, 8, 16, 32, 64 and 128 denoising steps. Lines average the 3 training seeds; bands show ±1 SD across seeds.

A successful global prediction has mean position error at most 0.1 world units; a successful local prediction has five-frame physics error at most 0.05. Success fractions are averaged across seeds, and the table gives the smallest tested budget at which at least 90% of clips pass. “None” means no tested budget reaches that.

BallsGenerated framesSteps for global successSteps for local successBest mean local error
12021280.0338
1282320.0297
1362320.0273
1442320.0256
220NoneNone0.1755
228NoneNone0.2525
236NoneNone0.2972
244NoneNone0.3353
320NoneNone0.3895
328NoneNone0.5398
336NoneNone0.6254
344NoneNone0.6293
520NoneNone0.6840
528NoneNone0.8318
536NoneNone0.8848
544NoneNone0.8927

The 1-ball, 20-frame cell needing 128 steps for the local threshold while 28, 36 and 44 frames need 32 runs opposite to a length-based account, and is unexplained. Download the sampling summary.

More denoising steps raise the score error the bound uses and lower the physics error

Download PNG
Two panels against denoising step count: the score error averaged over the noise levels a run of that many denoising steps evaluates the network at rises at every ball count, while the five-frame local position error of the clips those runs produce falls and then flattens
Figure 28. Left: \(\varepsilon_{\mathsf{score}}\), the squared score error averaged over the noise levels a run of \(K\) denoising steps evaluates the network at, over the 1,024 held-out clips and the 8 noise draws, then square-rooted. Right: the mean five-frame local position error of the clips those same runs produce, in world units, the same quantity as Figure 27, at 44 generated frames. Both panels use the 49-frame models; lines average the 3 training seeds and bands show ±1 SD across seeds. Both axes are logarithmic. The left panel starts at 2 steps because a one-step run evaluates the network only at \(\sigma=1\), where the score error is exactly zero.

The quantity the theorem assumes bounded grows with the denoising step count

The bound in section 01 takes one number from the model: \(\varepsilon_{\mathsf{score}}\), which Li and Yan define as the squared score error averaged over the denoising steps taken. The tables above average over a fixed grid of 32 uniform noise levels, which is not the grid any denoising run uses. A run of \(K\) steps evaluates the network at \(\sigma=1,(K-1)/K,\dots,1/K\). Averaging the audit’s saved per-level numbers over each of those grids gives \(\varepsilon_{\mathsf{score}}\) as a function of \(K\), with no further network evaluations.

The score error at level \(\sigma\) equals \((1-\sigma)/\sigma\) times the flow-prediction error, so it grows like \(1/\sigma\) as \(\sigma\) falls. The lowest level a \(K\)-step run visits is \(1/K\), and that single level contributes most of the average. So the average rises with the step count, at every ball count, the one-ball control included.

Balls2 steps48163264128
10.71.12.55.611.121.239.9
24.77.011.419.734.460.0103.4
36.69.515.225.944.676.2130.6
59.012.319.131.754.393.4160.7

Table: \(\varepsilon_{\mathsf{score}}\) on the noise levels each denoising run uses, 49-frame models, 44 generated frames, means over 3 training seeds. Each four-fold increase of the step count multiplies it by about 3. Dropping every second audited noise level and rebuilding it from its neighbours by the same interpolation reproduces the held-out squared error with a median error of 1.1% over all 48 models; the largest error is 8.7%, at the lowest of the 32 uniform levels, where dropping a level leaves a gap wider than any the table interpolates across. Source: experiments/score-error-audit/SCORE_VS_STEPS.md. Download the full grid, all four clip lengths and all three seeds.

Over the same range of step counts the five-frame local position error falls and then flattens. The bound’s score-error term rises while the accuracy of the clips those same runs produce improves, so the bound’s shape in the step count does not follow the measurement’s shape in the step count.

The measured score error is above the largest the bound allows at every step count but one

Download PNG
Two panels against denoising step count: on the left the first term of the bound rises to a largest value at 16 steps and then falls while the second term rises at every step count, on the right the measured score error rises above three dashed lines showing the largest error the bound allows
Figure 29. Left: the two terms of the bound in section 01, each divided by the theorem constant \(C\). The grey line is \((\log K)^3/K\), the first term divided by the discrete intrinsic dimension \(d_{\mathrm{int}}\). The coloured lines are \(\varepsilon_{\mathsf{score}}(K)\sqrt{\log K}\), the second term with the measured score error, at 1 and 5 balls. The dotted line marks 1, the largest a total variation distance can be. Right: the measured score error of Figure 28, with dashed lines for the largest score error that keeps the bound at or below 1 when \(C=1\), at three values of \(d_{\mathrm{int}}\). Both panels use the 49-frame models and average the 3 training seeds. All axes are logarithmic.

Two denoising steps minimise the bound at every intrinsic dimension up to a million

The theorem constant \(C\) and the discrete intrinsic dimension \(d_{\mathrm{int}}\) have no known values, so the bound has no numerical curve. The first term divided by \(d_{\mathrm{int}}\) is \((\log K)^3/K\), which needs no measurement, and the second term is \(\varepsilon_{\mathsf{score}}(K)\sqrt{\log K}\), which the measurement above supplies. Both can be drawn against \(K\).

Over \(K\) from 2 to 128 the first term rises from its 2-step value to its largest value at 16 steps and then falls, and it stays above its 2-step value until \(K\) is about 3,000. The second term rises at every step count. Minimising their sum over \(K\) from 2 to 1,024 therefore returns 2 steps for every \(d_{\mathrm{int}}\) from 1 to 1,000,000. Of those minima only the one-ball value, and only with a small intrinsic dimension, is below 1.

Setting both \(C\) and the bound to 1, the largest value a total variation distance can take, gives the largest score error the bound permits at each step count:

\[\varepsilon_{\max}(K)=\frac{1-d_{\mathrm{int}}(\log K)^3/K}{\sqrt{\log K}}.\]

A larger \(C\), or any accuracy target below 1, lowers that value. Both choices made here therefore permit a larger score error than any other choice would.

Balls2 steps48163264128
10.581.383.9910.6823.8048.8196.53
24.028.8818.5237.7873.67137.90250.14
35.5511.9824.6649.7395.39175.14315.86
57.5915.5530.9660.93116.23214.57388.77

Table: the measured score error divided by the largest the bound permits, at \(d_{\mathrm{int}}=0.1\). 49-frame models, means over 3 training seeds. One cell is below 1: one ball at 2 steps. Everywhere else the measured value is larger, and the ratio grows with the step count.

The average above runs over 32 uniform levels under a flow-matching convention rather than Li and Yan’s timestep grid, so it is not a numerical evaluation of their \(\varepsilon_{\mathsf{score}}\).

The datasets

Every clip comes from its own simulation: 400 frames simulated per seed, the first 200 discarded as burn-in, at most one accepted window kept. The original generator supplies the initial conditions, solver, window filters, collision-density limits and the held-out collision-count balancing; training and evaluation use disjoint simulation seeds. Each training set holds 20,000 clips from 20,000 simulations, each held-out set 1,024 from 1,024. Section 03 explains why this replaced the earlier datasets, which drew up to 1,000 overlapping clips from each of about twenty simulations.

BallsGenerated framesTraining minTraining meanHeld-out minHeld-out maxHeld-out mean
120–4400.00000.00
22022.17232.50
24444.18496.50
32022.69243.00
34455.685138.99
52046.53486.00
544813.9582214.98

Ball–ball collisions per clip, counted with the original generator’s frame-selection rule; the 28- and 36-frame rows fall between the two shown. The evaluation set balances the allowed collision counts. Download the full dataset validation record.

Numerical precision

Sampling runs in bf16, so one shared noise draw per level compares bf16 inference against float32 at 49 frames. Both ran on the same GPU, an RTX PRO 6000 Blackwell in one Slurm job, so the first ratio isolates the dtype. The second re-runs the same float32 configuration on an A100, isolating the hardware.

BallsSeedbf16 / float32, same GPUfloat32 on Blackwell / float32 on A100
101.02141.0000
201.01691.0000
301.01161.0049
511.02010.9988

bf16 raises the total squared score error by 1.2 to 2.2 percent; moving the same float32 computation to different hardware changes it by at most 0.5 percent. Neither moves any number above enough to change a conclusion.

How this relates to the two papers

What the proofs establish. For deterministic prediction the exact conditional score contains the correct future: one score evaluation plus arithmetic recovers it, which is Seriality Gap, Proposition 4.1. Serial Scaling extends this to approximate scores under further assumptions; its error bound must stay small enough across input sizes, which with the other conditions gives a solver of constant circuit depth (Appendix F). Appendix F defers the score-error assumption itself to Li and Yan (2024), where it reads

\[\varepsilon_{\mathsf{score}}^{2}:=\frac{1}{T}\sum_{t=1}^{T}\mathbb E\left[\left\|s_{t}(X_{t})-s_{t}^{\star}(X_{t})\right\|_{2}^{2}\right].\]

The squared L2 norm runs over all \(d\) output coordinates with no \(1/d\) normalization, and the average is over timesteps only. In their Theorem 1 it enters the bound as \(c\,\varepsilon_{\mathsf{score}}\sqrt{\log T}\), linearly and with no \(\sqrt d\) factor. That is why we report the unnormalized total alongside the per-component figure: the total is the quantity the assumption bounds. Our average runs over 32 uniform noise levels under a flow-matching convention rather than their timestep grid, so our number is not a numerical evaluation of their \(\varepsilon_{\mathsf{score}}\).

This also says when accurate scores cannot come from a shallow network: if a task needs growing circuit depth, a fixed-depth network cannot keep satisfying the conditions. The theory therefore allows failure to learn accurate scores on hard tasks, and does not require first observing an accurate-score model and then watching it fail. What it does not do is bound the computational limits of an arbitrary inaccurate denoiser. This audit measures score accuracy; it does not read circuit depth off an error curve.

What happens with one ball

One ball collides only with walls, so its future is a closed-form function of the given state with no chain of dependent events. It is the paper’s control for the claim that clip length alone is not what hurts one-pass generation.

1-ball measurement20 generated frames44 generated frames
Paper: local physics error0.2200.246
Paper: distance from true trajectory0.3101.417
Our models: local physics error0.0360.028
Our models: distance from true trajectory0.0240.044
Our models: score error per predicted number4.5754.426

The paper’s values come from its Table 2; its 25- and 49-frame clips include five conditioning frames. Our rows use the 25,000-update checkpoints, 64 denoising steps, and means across 3 training seeds.

Local physical consistency is flat, and slightly better on the longer clips. Score error per predicted number is flat. Distance from the true trajectory grows from 0.024 to 0.044 but stays two orders of magnitude below the box size: a marginally wrong velocity accumulates positional drift over a longer horizon even while the motion stays physically correct. The paper’s own 1-ball global error grows for the same reason, and by far more.

Two claims in the earlier state-diffusion write-up do not survive and are withdrawn. First, that single-ball wall bounces are a mild version of the same serial limitation: on these datasets block-causal generation beats one-pass generation by 1.16× to 1.31× at one ball, flat across the tested range, against 2.94× to 3.98× at five balls, and the paper’s two video models differ by 1.12× in the same comparison. Second, that the paper’s flat 1-ball curve sits on a VAE-and-tracker measurement floor of 0.19: no floor needs to be invoked, because the pixel models and the state models agree.

The total grows because there are more numbers to predict

From 20 to 44 generated frames the output grows by 2.2×, so a flat per-number error predicts √2.2 = 1.48× growth in the total. Measured: 1.46× at one ball, 1.48× at two, 1.50× at three, 1.37× at five. A total that grows as the square root of output size, with no increase per number, is what a model with a fixed positive per-number error produces on a larger output: if that error settles at any constant \(a\), the total is \(\sqrt{4BF}\,a\), which grows without limit as \(F\) increases.

Holding the total bounded would require the per-number error to shrink like \(1/\sqrt d\). Over a 2.2× growth in output, \(\sqrt d\) grows by 1.48, so the per-number error would have to fall to 0.67× to keep the total fixed. Measured instead:

BallsPer number, 20 framesPer number, 44 framesObserved changeNeeded for a fixed total
12.1382.1040.98×0.67×
24.1514.1361.00×0.67×
34.2324.2721.01×0.67×
54.3844.0580.93×0.67×

No ball count comes close, and the one-ball case is the most informative: there the task has a closed-form solution, the model learned it — physical accuracy is flat and slightly improving with length — and the per-number score error is flat. That is about as favourable as this setting gets, and the total still grows as √d, because flat per-number error and a growing output force it to.

So the growing total is not evidence that score accuracy breaks down as the requested future lengthens. It shows that the assumption, stated as a bound on an unnormalized sum over a growing number of coordinates, would require a model to become more accurate at each individual number as more numbers are asked of it. We see no sign of that and no reason to expect it. A separate structural point runs the same way: the first term of Li and Yan’s bound, \(c\,d\log^{3}T/T\), carries \(d\) whatever the score error is, so even an exact score would need the step count \(T\) to grow roughly linearly in \(d\) to hold that term at a fixed target. We have not examined how Appendix F treats this.

What four frame counts can and cannot show

The assumption is not that the score error is small in this experiment. It is that one constant works for every input size: there exists some \(C\), fixed once and for all, with \(\varepsilon_{\mathsf{score}}(d)\le C\) for every output size \(d\). That is a claim about infinitely many sizes. We measured four.

Four measurements cannot show the assumption fails. A sequence can rise and still be bounded — it might climb towards a ceiling and never pass it. These totals could level off at 60 generated frames, or at 600, and the assumption would hold.

They also cannot show it holds. Four numbers have a largest value, so they are trivially bounded by it, and “bounded over the range we tested” is true of any finite set of measurements.

The question becomes testable through the exact identity between the two quantities in this section. Writing \(a(d)\) for the error per state component, \(\varepsilon_{\mathsf{score}}(d)=\sqrt d\,a(d)\), so

\[\varepsilon_{\mathsf{score}}(d)\le C \quad\Longleftrightarrow\quad a(d)\;\le\;\frac{C}{\sqrt d}.\]

The assumption is therefore exactly the per-component error shrinking at least as fast as \(1/\sqrt d\) — the same claim rewritten, not a weaker version — and that is a trend we can measure. What we measure is that \(a(d)\) does not shrink at all: over a 2.2-fold increase in \(d\) it stays within 7% of where it started, at every ball count, where a 33% drop would have been needed.

The totals say the same from the other side. Fitting \(\log\varepsilon_{\mathsf{score}}=\text{const}+p\log F\):

BallsTotals at 20, 28, 36, 44 generated framesExponent \(p\)
119.1, 23.1, 25.9, 27.90.482
252.5, 62.8, 71.4, 77.60.499
365.6, 81.0, 91.7, 98.20.518
587.7, 103.3, 112.2, 120.40.399

The largest residual is 0.021 in log space. The exponents sit at 0.40 to 0.52, so over this range the growth is the pure √d rate that flat per-component error forces, with no sign of flattening. The successive slopes drift down slightly, and that is accounted for by the small decline in per-component error — the 5-ball models have both the lowest exponent and the largest per-component decline — rather than by approaching a ceiling.

The honest conclusion. We have not refuted the assumption and do not claim to. What we can say is sharper than “unsupported”: the assumption is equivalent to a specific measurable trend, we measured that trend, and it is absent. For the assumption to hold, per-component accuracy would have to start improving as the model is asked for more output, somewhere past 44 generated frames, having been flat throughout the range we tested. We know of no mechanism that would produce that. One further constraint is easy to lose sight of: the theorem needs \(\varepsilon_{\mathsf{score}}\) not merely bounded but small enough for the accuracy target, so a large constant would not rescue the argument. Boundedness is the weaker of the two requirements, and it is the one already not in evidence.

Very low noise dominates the average

Converting a predicted flow to a score multiplies the flow error by \((1-\sigma)/\sigma\), which is large when \(\sigma\) is small. The lowest primary level, \(\sigma=1/64\), contributes 94 to 97 percent of the total squared score error at one ball and 91 to 92 percent at two, three and five balls, at every frame count. Because of that weighting, the headline averages could hide behaviour at moderate noise, so we recompute per predicted number over \(\sigma\ge0.1\) only:

Balls20 generated frames44 generated framesChange
10.12420.16411.32×
20.60440.55660.92×
30.63450.60650.96×
50.64480.59370.92×

At moderate and high noise the multi-ball models are flat or slightly improving, and the 1-ball models show a residual 1.32× rise with length. Two things about that number matter. It is small in absolute terms — the 1-ball error stays four times below every multi-ball value, so the model that grows is the accurate one. And it is far from the earlier models’ behaviour, where the same restricted measurement rose by roughly 24× across the same range. We report it rather than average it away, and we have no explanation for it.

Why score error and generation error need not move together

With \(z_\sigma=(1-\sigma)y+\sigma\epsilon\) and \(y\) the true standardized future, the clean-state estimate from the same network call is \(\hat y=z_\sigma-\sigma\hat v_\sigma\), and its error relates to the score error by

\[\hat s_\sigma-s^*_\sigma=-\frac{1-\sigma}{\sigma}\left(\hat v_\sigma-v^*_\sigma\right),\qquad \hat y-y=\frac{\sigma^2}{1-\sigma}\left(\hat s_\sigma-s^*_\sigma\right).\]

For \(0<\sigma<1\) a large score error at low noise can correspond to a small error in the reconstructed clean state. Reconstruction from a noisy true future is a denoising measurement; ordinary generation starts from pure noise plus the conditioning frames. The two answer different questions, so aggregate score error and generation error need not change in proportion.

The pattern is reproduced; the mechanism is not demonstrated

The Seriality Gap paper’s central evidence is a pattern across clip lengths: with five balls, one-pass generation gets worse as clips get longer while generating in order does not, and with one ball neither gets worse and the two are close. We reproduce the pattern. Five-frame physics error, 25-frame clips to 49-frame clips:

ModelPaperThis work
5 balls, one-pass0.670 to 0.8830.684 to 0.885
5 balls, in order0.613 to 0.6070.233 to 0.223
1 ball, one-pass0.220 to 0.2460.037 to 0.030
1 ball, in order0.220 at 49 frames0.032 to 0.023

What the audit adds is that these models do not satisfy the assumption the depth argument rests on. The theorem that limits one-pass diffusion to constant depth requires an accurate score. Ours is not accurate at any ball count, by the measurements above. So the degradation we observe falls outside the theorem’s reach and cannot be attributed to it. A plainer explanation covers the same observations: the model does not learn an accurate score for long multi-ball futures. That is a learning failure, and in these metrics it is indistinguishable from a depth barrier.

Across the grid, score accuracy and generation quality coincide exactly. One ball is the only case where the score error per state component is small, roughly a quarter of the multi-ball values, and it is also the only case where quality holds up with length and where the two architectures are close. Every cell that degrades is a cell whose score is poor.

The theory itself allows this, which is why the ambiguity is hard to remove. A fixed-depth network asked for a task needing growing depth should fail to learn accurate scores, so a rising score error is what the theory predicts on such a task. But a model that is too small, trained too briefly, or facing a task that is merely hard also has a rising score error. Score error alone cannot tell them apart, and neither can generation quality, because both explanations predict the same degradation. What would separate them is a multi-ball setting where the score is known to be accurate, so a failure could not be blamed on score quality. Two routes are in the open questions below.

What these experiments establish

At one ball, on a task with a closed-form solution, these models keep both their physical accuracy and their per-number score accuracy as the requested future grows, and one-pass generation is nearly as good as generating in order. At two or more balls, physical accuracy degrades with length and generating in order is markedly better, while per-number score accuracy stays flat. No tested denoising-step budget up to 128 fixes the multi-ball case. No model at any ball count satisfies the low-score-error assumption: the total rises with output size in every case, and the per-component decay that boundedness requires is absent.

What this does not establish: that score error is unbounded as length grows without limit, since four lengths were tested; that the models satisfy the remaining conditions of the approximate-score theorem, which we do not measure; that any observed error curve implies a particular circuit depth; or that the multi-ball degradation is caused by a depth limitation rather than by the model failing to learn an accurate score. The last is the main interpretive limit of this audit.

Open questions

  • The residual 1-ball rise at moderate noise. 1.32× across the tested range, with no explanation. Longer training, wider models, or comparison against an analytic single-ball predictor and its exact score would show whether it is an optimization artifact or a property of the task.
  • A multi-ball setting with a known-accurate score. Training against an analytically computed score, so a failure to generate well cannot be attributed to score quality. This is the direct way to separate a depth limitation from a learning failure.
  • Interacting against independent balls at identical output size. Same ball counts, frame counts and initial-state distribution, with ball–ball collisions enabled or disabled and wall bounces kept. Without ball–ball collisions each ball has an independent closed-form solution, so this separates the cost of predicting more numbers from the cost of the balls affecting one another. Not yet run.

Scope and data

What was run

48 score audits — 4 ball counts × 4 frame counts × 3 seeds — and 384 physics evaluations, all complete. Each audit is 1,024 held-out clips × 32 primary noise levels × 8 noise draws, one network evaluation per noisy input.

Tables behind the figures

The per-seed score and sampling numbers, the per-clip spread statistics, and the dataset validation record.

04 / State-space guidance

Guidance using the cleaner
state diffusion set-up

CollapseExpand

One gradient step toward a physics target at every denoising step. In state space the guides that failed in pixels now work — they lower every error of the bidirectional model — but they shift the whole length curve down rather than flattening it. Guidance built from physics laws alone, plus renoise-and-denoise, does flatten it, and still lands above the block-4 model.

5 balls · 25–49 frames 896 holdout clips Seeds 0–2 1,121 guided runs

Every number here is on the 896 holdout clips — the 1,024 held-out clips minus 128 kept for calibration — as the mean ± standard deviation over 3 training seeds, using the retrained bidirectional state models from the previous section at 25,000 updates. Guidance strength is selected per (frames, steps, guide) on the calibration clips; the selected value is 1 in every main-grid cell except three 20-step cells, where it is 0.3. Source: experiments/statebench-guidance/REPORT.md.

What we found

  1. The rollout targets work here and failed in pixels. The same targets changed pixel-space error by within ±2%, because a tracker had to recover ball positions from video first. On exact states there is nothing to detect: rollout-5 guidance lowers the five-step local error by 28% at 49 frames and 34% at 25, and improves 84–87% of individual clips.
  2. They do not flatten the length curve. The 25→49-frame growth of the local error is 1.28× unguided, 1.26× with a rollout-1 target and 1.41× with rollout-5. Guidance moves the whole curve down by roughly a constant factor, and the guided model stays about three times worse than block-4.
  3. Rollout guidance fixes positions and breaks velocities. Overlapping frames halve and the collision-count distribution is repaired, but energy drift triples and by frame 48 a rollout-5 sample has kept 47% of its initial kinetic energy. Neither local error catches this; a conserved quantity tracked across the whole clip does.
  4. Guidance from physics laws alone draws with block 4 when both get the renoise-and-denoise rounds, and loses to it when each uses its own best setting. Those rounds are the reason. They help every method except block 4, which they damage, so forcing both onto them flatters the law set. Run block 4 without them and it leads on both local errors at every clip length.
  5. The law set's error stops growing with clip length, but only with the rounds. Take them away and it grows like every other guided method. The flat curve needs the laws and the rounds together, and neither produces it alone.

Error against clip length

Download PNG
Three panels of error against clip length at 50 denoising steps: vanilla, rollout-1, rollout-5 and oracle guidance, with the unguided block-4 model as a dotted reference
Figure 30. 50 denoising steps. Dotted orange is the block-4 model without guidance, taken from the state-diffusion evaluations on all 1,024 clips. Guidance lowers every curve and leaves every slope alone.
Δx(5), local error25f33f41f49f25→49
bidirectional, vanilla0.691 ±0.0110.838 ±0.0250.880 ±0.0060.888 ±0.0141.28×
bidirectional, rollout-1 target0.550 ±0.0050.649 ±0.0230.687 ±0.0110.693 ±0.0171.26×
bidirectional, rollout-5 target0.453 ±0.0080.550 ±0.0170.609 ±0.0080.639 ±0.0091.41×
bidirectional, oracle0.147 ±0.0050.170 ±0.0070.181 ±0.0060.188 ±0.0021.28×
block-4, vanilla0.233 ±0.0040.239 ±0.0020.233 ±0.0040.223 ±0.0020.96×
Δx(1), local error25f33f41f49f25→49
bidirectional, vanilla0.143 ±0.0030.158 ±0.0060.160 ±0.0020.159 ±0.0041.11×
bidirectional, rollout-1 target0.099 ±0.0010.111 ±0.0040.115 ±0.0020.116 ±0.0041.18×
bidirectional, rollout-5 target0.093 ±0.0020.106 ±0.0040.114 ±0.0020.119 ±0.0011.28×
bidirectional, oracle0.026 ±0.0010.030 ±0.0010.031 ±0.0010.033 ±0.0001.28×
block-4, vanilla0.036 ±0.0010.037 ±0.0000.035 ±0.0010.034 ±0.0010.93×
Δx(GT), global error25f33f41f49f25→49
bidirectional, vanilla1.325 ±0.0082.098 ±0.0142.602 ±0.0142.934 ±0.0122.21×
bidirectional, rollout-1 target1.218 ±0.0051.965 ±0.0262.492 ±0.0062.836 ±0.0102.33×
bidirectional, rollout-5 target1.056 ±0.0131.795 ±0.0232.343 ±0.0102.704 ±0.0132.56×
bidirectional, oracle0.032 ±0.0010.034 ±0.0010.038 ±0.0020.042 ±0.0011.31×
block-4, vanilla0.792 ±0.0081.455 ±0.0081.988 ±0.0132.358 ±0.0172.98×

The rollout targets work in state space

In pixel space the same targets gave changes within ±2%, which the rollout-guidance section above documents. Here rollout-5 guidance lowers Δx(5) by 28% at 49 frames and 34% at 25 frames, and improves 84–87% of individual holdout clips when paired on the same initial noise. The pixel-space failure was a tracking problem, not a property of the guidance idea: with exact states there is nothing to detect.

The rollout targets do not flatten the length curve

The 25→49 growth of Δx(5) is 1.28× unguided, 1.26× with rollout-1 and 1.41× with rollout-5. The growth of the global error is unchanged or slightly steeper, 2.21× against 2.33–2.56×. Guidance shifts the whole curve down by a roughly constant factor. The block-4 model is flat at 0.96× and still 2.9× more accurate than the best guided bidirectional model at 49 frames.

A rollout-5 target improves Δx(1) as much as a rollout-1 target does

0.119 against 0.116 at 49 frames, and it improves Δx(5) and Δx(GT) more. The five-frame target carries more information about where the ball should be than the one-frame target, whose source frame is itself only one frame away from the target.

The oracle

The oracle receives the true future states. It is a positive control: it shows how far one gradient step per denoising step can move the sample. At 50 steps it takes the global error from 2.93 to 0.04 and the five-step local error from 0.888 to 0.188, below the block-4 model's 0.223. It also gains from more steps, which vanilla sampling does not:

Δx(5) at 49 frames20 steps50 steps100 steps
vanilla0.9000.8880.886
rollout-1 target0.7660.6930.687
rollout-5 target0.6990.6390.622
oracle0.3900.1880.123

At 100 steps the oracle's 25→49 growth is 1.08×. This matches the pixel-space clean-latent oracle, which also kept improving with steps while vanilla sampling did not. The guided steps are an iterative fit to a supplied answer; they do not compute the answer.

Error against guidance strength for the three targets at 49 frames and 50 steps, with a minimum near strength 1 and every guide worsening by strength 10
Figure 31. 49 frames, 50 steps. Strength 1 makes the correction the size of a typical Euler update.

Below 0.3 the rollout guides do little; at 3 they start to hurt Δx(1); at 10 every guide is worse than no guidance on the local errors, the oracle included. The optimum is 1 for all three targets, and the same value is selected at every clip length and step count from 50 steps upward.

Ablations

One change at a time from the main grid, 49 frames and 50 steps, with strength re-selected on the calibration clips.

Variantrollout-1: Δx(1) / Δx(5)rollout-5: Δx(1) / Δx(5)oracle: Δx(5) / Δx(GT)
main grid0.116 / 0.6930.119 / 0.6390.188 / 0.042
source velocities from the paper's estimator instead of the model's0.113 / 0.6610.120 / 0.611—
positions only in the loss0.113 / 0.7350.133 / 0.7440.448 / 0.090
guidance only in the second half of the steps0.140 / 0.8070.141 / 0.8060.427 / 0.180
gradient without the model Jacobian0.118 / 0.7510.149 / 0.7680.122 / 0.027
  • Re-estimating the source velocities from positions, as the metric does, gives the best rollout results — a further 4–5% on Δx(5). The model's own velocities are slightly less consistent with its positions than finite differences are.
  • Dropping the velocity channels from the loss hurts every guide: the velocity targets carry information the position targets do not.
  • Guiding only late loses most of the gain. The early, high-noise steps matter even though the clean estimate there is a blurred mean trajectory.
  • For the rollout targets the gradient through the model helps; for the oracle it hurts, and the plain Tweedie-estimate gradient gives the best oracle numbers of the whole experiment, Δx(5) 0.122 at 50 steps. With a fixed exact target there is no need to ask the model which direction is consistent with its own prediction.

What the methods actually produce

Figure 32. Five holdout clips drawn at random, 49 frames, training seed 0, 50 steps. Panels left to right: ground truth, vanilla, rollout-1 guided, rollout-5 guided, oracle guided, block 4. Every generated panel starts from the same initial noise, so the differences between them are due to the guidance alone; block 4 uses its own noise. Frames 0 to 4 are the given frames and are identical in every panel. Each ball drags a fading trail of its previous 8 frames, the same rendering as the state-diffusion gallery. Each panel carries that clip’s own Δx(5) and Δx(GT).

Watching the trails shows what the metrics measure. The vanilla panel produces balls that turn in mid-flight, pass through each other and leave the wall at the wrong angle. Rollout-5 guidance removes many of these, so its trails look like a plausible billiard trajectory, but by the end of the clip the balls are nowhere near their true positions: a physically consistent future that starts from slightly different collisions diverges from the true one within about 20 frames. The oracle panel follows the truth panel almost exactly. Block 4 produces a consistent trajectory that is also a different one.

Clip 112 as still frames

Download PNG
Grid of six frames by eight methods: each generated ball is a filled disc joined by a line to a dashed circle at its true position, and the lines grow with frame number for every method except the oracle
Figure 33. Holdout clip 112, the median vanilla Δx(5), with the two physics-law methods included as the last two rows. The filled disc is the generated ball; the dashed circle in the same colour is where that ball truly is at that frame, and the line joins the two — so the length of each line is that ball's contribution to the global error.

At frame 5 every method sits on the truth. By frame 20 the vanilla and both guided runs have visibly separated, and by frame 36 the lines are as long as the box. Only the oracle keeps every disc inside its dashed circle, which is why it has no visible lines. Guidance shortens these lines a little but does not stop them growing.

Distance to the truth, frame by frame

Download PNG
Mean distance to the true position against frame number: vanilla and both guided curves rise to the same plateau, block 4 rises later, the oracle stays flat and near zero
Figure 34. The same quantity as Figure 25, averaged over the 896 holdout clips. Guidance with rollout targets delays the divergence by a few frames, block 4 delays it more, and all three reach the same plateau of unrelated positions by frame 40. Only the oracle stays near the truth.

More plausibility measures

The paper's two metrics compare two points in time. These ask what the simulator says about a generated trajectory as a whole, on the same trajectories as Figure 22 at 50 steps and strength 1. They fall into three groups: what the model's own velocity channel says, what the simulator says about the generated positions, and how far the sample is from the truth.

“Event” everywhere means what the paper's simulator reports when run for one frame interval from the generated frame t−1 with the generated positions and velocities. Because the true clips are exact simulator snapshots, the ground truth scores zero on every violation measure, which checks the definitions.

Every measure, every method, both clip lengths

Download PNG
Dashboard of small multiples: each panel is one plausibility measure against clip length, with a line per method
Figure 35. The figures in this group also carry the two physics-law rows described below, which were added later. Along clip length, every velocity-based measure of the guided models gets worse from 25 to 49 frames while block 4 and the oracle stay flat. The position-based measures are flat for everyone except vanilla's collision count, which grows from 4 to 10 excess collisions per clip.
49 framesground truthvanillarollout-1rollout-5oracleblock 4
energy drift0.0000.057 ±0.0020.125 ±0.0020.172 ±0.0040.020 ±0.0010.019 ±0.000
speed change without ball contact0.0000.030 ±0.0020.044 ±0.0020.054 ±0.0010.011 ±0.0010.007 ±0.000
collision-law residual, ball–ball0.0000.706 ±0.0060.428 ±0.0080.576 ±0.0350.130 ±0.0030.229 ±0.007
collision-law residual, wall0.0000.213 ±0.0040.234 ±0.0080.245 ±0.0060.052 ±0.0040.046 ±0.003
phantom velocity-change rate0.0000.067 ±0.0040.105 ±0.0050.136 ±0.0070.017 ±0.0000.014 ±0.000
frames with overlapping balls0.0000.140 ±0.0110.079 ±0.0110.071 ±0.0040.004 ±0.0010.008 ±0.001
frames with a ball outside the box0.0000.032 ±0.0010.015 ±0.0000.031 ±0.0030.004 ±0.0000.002 ±0.000
position–velocity mismatch (units)0.0000.019 ±0.0010.030 ±0.0010.036 ±0.0010.011 ±0.0000.008 ±0.000
ball–ball collisions per clip15.125.018.915.415.114.7
W1 distance, collision counts0.0009.830 ±0.7483.829 ±0.9771.475 ±0.2090.185 ±0.0630.544 ±0.006
W1 distance, speeds (units/s)0.0000.231 ±0.0041.345 ±0.0431.715 ±0.0340.060 ±0.0010.109 ±0.016
median frame of divergence from truth491314154918

896 holdout clips, mean ± sd over 3 seeds. The Wasserstein-1 distance is the mean absolute difference between the quantile functions of the generated and true distributions.

Rollout guidance fixes positions and breaks velocities

On the position-based measures the guided models sit between vanilla and block 4: half as many overlapping frames, the ball–ball collision-law residual down from 0.71 to 0.43 with a rollout-1 target, and the collision-count distribution repaired — vanilla generates 25 collisions per clip against 15 in the data, and rollout-5 gives 15.4. On every velocity-based measure the guided models are worse than vanilla: energy drift 3×, phantom velocity changes 2×, speed-distribution distance 6×.

The velocity channel loses energy through the clip

Download PNG
Fraction of initial kinetic energy against frame number: both rollout-guided curves fall steadily while vanilla, oracle and block 4 stay near one
Figure 36. By frame 48 the rollout-5 sample keeps 47% of its initial kinetic energy, rollout-1 65%, vanilla 95%, the oracle and block 4 98%.

The rollout target's velocity at frame t is the simulator's velocity after running from the estimate's frame t−k. Early in sampling the estimate is a conditional mean whose velocities are averaged over many futures and therefore too small, and the target inherits and then enforces those small velocities. Neither local error catches this: the paper's re-estimates velocities from positions, and the generated-velocity variant only checks that a few frames are consistent with each other, which a trajectory that is too slow but self-consistent passes — both variants improve under rollout guidance. A conserved quantity tracked across the whole clip is what exposes it.

The oracle and block 4 are physically clean by every measure; vanilla is not

Vanilla's ball–ball collision-law residual is 0.71, meaning the velocity after a detected contact differs from the elastic-collision outcome by 71% of the incoming speed, against 0.13 for the oracle and 0.23 for block 4. Vanilla balls overlap in 14% of generated frames and leave the box in 3%; the oracle and block 4 do so in under 1%.

Collision counts and speeds

Download PNG
Distributions of ball-ball collisions per clip and of ball speeds, one curve per method against the true distribution
Figure 37. Vanilla collides far too often; block 4 and the oracle match the data.

Guidance delays divergence by one or two frames; block 4 by five

Download PNG
Distribution of the frame at which a clip first exceeds one ball radius from the truth, per method
Figure 38. Half of vanilla's clips are more than a ball radius from the truth by frame 13, guided clips by frame 14 to 15, block 4 by frame 18. The oracle never diverges.

Rollout error at every horizon

Download PNG
Local error against rollout horizon k, one curve per method, keeping the same order at every horizon
Figure 39. The curves keep their order at every horizon, so the choice of k = 1 and k = 5 does not change any conclusion, and the guided models track vanilla's shape rather than block 4's.

Are the generated clips distributed like true clips?

The measures above pick one physical quantity each. The complementary question is whether the generated clips, taken as whole objects, are distributed like true clips. Each clip's 44 generated frames are flattened into one vector — 440 position coordinates, 440 velocity coordinates, or both — in the model's standardized units, and compared with the truth three ways.

  • Gaussian Fréchet distance, the FID construction, between the 896 generated clips and 20,000 independent true clips of the same length. Fitting a Gaussian to 896 samples in 880 dimensions is noisy, so the ground-truth column — the 896 holdout true clips scored the same way — is the noise floor.
  • Sliced Wasserstein-2 distance to the same 20,000 clips: the exact 1-D W2 along 1,000 random directions, root-mean-squared. No Gaussian assumption.
  • Classifier two-sample test: 5-fold cross-validated accuracy of a classifier trained to tell the 896 generated clips from the 896 true clips they continue, where 50% means indistinguishable. A generated clip and its true continuation always share a fold; otherwise the twin in the training fold teaches the opposite label and accuracy falls below chance.
49 framesground truthvanillarollout-1rollout-5oracleblock 4
Fréchet distance, positions2.553 ±0.0002.622 ±0.0294.060 ±0.0304.105 ±0.0292.554 ±0.0012.589 ±0.017
Fréchet distance, velocities5.456 ±0.0005.420 ±0.0187.489 ±0.0207.740 ±0.0915.426 ±0.0045.420 ±0.019
Fréchet distance, both6.539 ±0.0006.558 ±0.0318.899 ±0.0299.147 ±0.0936.511 ±0.0026.526 ±0.030
Fréchet distance, both, top-32 PCs2.187 ±0.0002.286 ±0.0304.237 ±0.0344.160 ±0.0492.190 ±0.0022.239 ±0.045
sliced W2, positions0.0581 ±0.00020.0592 ±0.00170.0663 ±0.00130.0680 ±0.00230.0582 ±0.00030.0596 ±0.0005
sliced W2, velocities0.0638 ±0.00060.0664 ±0.00040.1809 ±0.00340.2161 ±0.00290.0637 ±0.00050.0610 ±0.0009
sliced W2, both0.0610 ±0.00040.0629 ±0.00030.1046 ±0.00170.1215 ±0.00250.0611 ±0.00040.0603 ±0.0006
classifier accuracy, temporal CNN0.505 ±0.0160.499 ±0.0030.930 ±0.0060.967 ±0.0020.508 ±0.0020.504 ±0.006
classifier accuracy, gradient boosting0.514 ±0.0040.545 ±0.0100.931 ±0.0030.967 ±0.0040.512 ±0.0030.509 ±0.004
classifier accuracy, logistic regression0.511 ±0.0060.500 ±0.0110.542 ±0.0160.534 ±0.0050.505 ±0.0030.496 ±0.007

Mean ± sd over 3 seeds, 49 frames. Distances in standardized units: positions divided by 2.5, velocities by 6.25.

Vanilla, the oracle and block 4 are indistinguishable from the truth on every whole-trajectory measure

Their Fréchet and sliced-W2 distances sit on the noise floor at every number of principal components, and the CNN and logistic regression score 50% against them. Gradient boosting reaches 54.5% on vanilla, the only trace of a signal. This is the same vanilla model that overlaps balls in 14% of its frames and collides 25 times per clip instead of 15. Those violations live in a few coordinates per frame — the distance between one pair of balls, the velocity jump at one contact — and change the marginal distribution of the 880 coordinates so little that a two-sample test on 896 clips does not see them.

Fréchet distance against the number of principal components

Download PNG
Fréchet distance against the number of principal components kept: vanilla, oracle and block 4 sit on the ground-truth noise floor at every dimension, while both rollout-guided methods are well above it
Figure 40. The conclusion does not depend on the dimension: only the rollout-guided models separate from the floor.

Only the rollout-guided models separate, and only through their velocities

The CNN tells rollout-1 and rollout-5 clips from true ones with 93% and 97% accuracy; sliced W2 on velocities is 3× the floor and on positions 1.15×. Logistic regression stays near 54%, so the difference is not a shift in the mean trajectory but in its shape.

What the classifier looks at

Download PNG
Classifier input sensitivity by channel and frame, concentrated on the velocity channels at frames 8 to 15
Figure 41. The sensitivity is concentrated on the velocity channels at frames 8 to 15, where the guided models lose kinetic energy fastest — from 97% to 74% of the initial energy between frames 5 and 15.

The separation grows with clip length

Download PNG
Velocity distance and classifier accuracy against clip length: both rise for the guided models while the other three methods stay on the noise floor
Figure 42. The guided models' velocity distance and classifier accuracy grow with every step from 25 to 49 frames while the other three methods stay on the floor.

Physics-feature space against raw coordinates

Download PNG
UMAP embeddings: in physics-feature space the methods form separate islands, while in raw-coordinate space they are one featureless cloud
Figure 43. Top row: each clip embedded by eleven per-clip physics measures from the table above, nothing relative to the truth. Bottom row: the raw 880 coordinates.

In physics-feature space there are three islands: the truth alone, since every violation measure is exactly zero for it; the two rollout-guided models, with their low final energy; and a third island where vanilla forms its own blob next to a shared oracle-and-block-4 blob. In raw-coordinate space every method is the same featureless cloud. That matches the table: the raw coordinates do not carry the physics violations that the simulator-based measures detect.

The whole-trajectory distances therefore add one thing — a clean confirmation that the rollout-guided velocity drift is a distribution-level defect — and miss everything else. Comparing generated trajectories as a distribution needs a feature space in which the violations are large, and for this task those features are the physical measures, not the coordinates.

Physics-law guidance

The guidance above pushes the model toward a target it computes for itself by simulating forward. That lowers the error, but the error still grows as clips get longer. This part drops the target and guides with physics instead.

A law is an equation relating neighbouring frames that the true motion satisfies exactly: energy is conserved, a ball touching nothing travels in a straight line, a ball leaving a wall reflects, two balls never overlap, and no ball passes through another. The model is pushed toward satisfying these. Nothing simulates the motion forward, so this guidance never needs to know what happens next — only what any correct answer must look like.

Three law sets were tried, and they are named below by what they contain rather than by the order they were tried in.

The law methods do more sampling work than what they are compared against

Everything earlier in this section runs the denoising pass once, applies its correction once per step, and produces one clip. The law methods add three things: renoise-and-denoise rounds, each mixing the finished clip back to a chosen noise level and running the steps below it again; a second model run per step, so the correction is applied twice; and repeated independent samples, where several clips are generated and one is kept. Each of the three lowers the error on its own, whatever the guidance is, so law results cannot be read against the earlier ones directly. This is what each method runs:

MethodRenoise-and-denoise roundsModel runs per denoising stepIndependent samples per clip (best kept)Total model runs per clip
vanilla, rollout-1, rollout-5, oracle; block 401150
the same four with renoise-and-denoise rounds (dashed in Figure 40)16 from σ 0.3511322
laws with a simulator16 from σ 0.3521644
rollout-5 given the same three settings, 49 frames only24 from σ 0.45184,624
geometry laws24 from σ 0.45289,248

The last column multiplies the other three together, so it counts how many times the model runs to produce one clip. The law methods cost far more than the methods they were originally compared against, which is why the comparison further down gives every method the same three settings.

49 frames, 50 steps, holdout, 3 seedsΔx(1)Δx(5)Δx(GT)
vanilla0.159 ±0.0040.888 ±0.0142.934
rollout-1 target0.116 ±0.0040.693 ±0.0172.836
rollout-5 target0.119 ±0.0010.639 ±0.0082.704
oracle0.033 ±0.0000.188 ±0.0020.042
rollout-5 target, 16 renoise-and-denoise rounds0.099 ±0.0030.531 ±0.0112.548
oracle, 16 renoise-and-denoise rounds0.018 ±0.0000.106 ±0.0020.020
block 4, vanilla0.034 ±0.0010.223 ±0.0022.358
laws + differentiable rollouts 1–3, 16 renoise-and-denoise rounds0.049 ±0.0020.317 ±0.0123.008

Three things made this work. Law gradients touch only a few numbers per frame, so each one needs its own size limit rather than a shared one — before that, every law made the result worse. Letting the gradient flow through a differentiable simulator beats comparing against a fixed simulated target, because it can also move the frame the simulation starts from. And renoise-and-denoise rounds help.

Renoise-and-denoise rounds help every method, though, not just this one, and they help the unguided model most of all, so they are not evidence that the laws are doing the work. Two further problems: this method still simulates forward, which is what the laws were meant to avoid, and its samples collide far less often than real ones.

How much of that came from renoise-and-denoise rounds alone

Download PNG
Local error against the number of renoise-and-denoise rounds, one curve per method. Every curve falls; the law curve is lowest but gains less from the renoise-and-denoise rounds than the unguided model does
Figure 45. Number of renoise-and-denoise rounds against error, on the longest clips. The law curve here is the first law set, the one that uses a differentiable simulator; the later law sets are not plotted. Unlike the comparison below, every series here is matched: one correction per step and one clip per attempt. Renoise-and-denoise rounds lower every method’s error. The laws are second on both measures — the unguided model gains the most in absolute terms and the oracle the most in relative terms — so renoise-and-denoise rounds are not a law-specific advantage. The hollow marker is the one point that re-noises to a different level, so it does not belong to its own curve, and no other method was run there.

Second law set: geometry laws alone, no simulator

This set drops the simulator entirely and keeps only the geometry laws: no overlaps, no crossings, stay inside the box. Every law the first attempt had rejected was retried with renoise-and-denoise rounds, five new laws were written, and the update rule was varied.

49 frames, 50 steps, holdout, 3 seedsΔx(1)Δx(5)Δx(GT)final energy
geometry laws, 4 renoise-and-denoise rounds0.0820.5682.9390.97
geometry laws, 16 renoise-and-denoise rounds, 2 inner corrections0.0710.5052.9450.97
geometry laws, 48 renoise-and-denoise rounds, 2 inner corrections0.0640.4632.9400.96
geometry + faded late reflection, 24 renoise-and-denoise rounds0.0660.4582.9440.96
the same, best of 8 candidates by the law critic0.055 ±0.0000.388 ±0.0022.9410.95
perfect pick of the same 8 (needs the simulator)0.0460.3152.913—
rollout-5 target, 24 renoise-and-denoise rounds, best of 8 by the same critic0.082 ±0.0020.449 ±0.0092.536—
oracle, 16 renoise-and-denoise rounds0.0180.1060.020—

Measured against methods run without renoise-and-denoise or repeated independent samples, this set looked like a clear win. It is not: give the rollout target the same renoise-and-denoise rounds and the same repeated independent samples and most of the apparent win disappears. The two figures below give every method the same settings, and differ only in whether each method keeps the best of its samples.

This set has three parts, and only the first is guidance:

  1. Geometry guidance. No overlaps, no crossings, inside the box; bounded per-element step, strength 3, plus the relative-velocity reflection law at weight 0.3, faded as it is satisfied, applied only when σ ≤ 0.35 and added without renormalising the geometry step. The reflection helper is worth 0.02; nothing else helps as a guide.
  2. Renoise-and-denoise, 24 rounds from σ 0.45 with two corrections per step: 0.654 → 0.458.
  3. Law-critic selection. Eight samples per clip from different noise; the one with the smallest sum of z-scored log losses of three laws — position-only free flight, speed constancy outside contacts, total momentum outside wall contacts, all fixed in advance on a separate run — is kept: 0.457 → 0.388. The critic captures about half of what a perfect pick would, 0.311.

Some laws help as a guide; others only help as a filter

Laws written on velocities get satisfied in the wrong way. The model has a velocity channel that the error metric never reads, so it can satisfy a velocity law by editing those numbers, and it can satisfy a collision law by producing fewer collisions. Both make the law's own residual look better while the motion gets worse.

Laws written on positions have the opposite problem: they need a recognisable contact to act on, and a badly wrong sample does not contain one, so they go quiet exactly where they are needed. A straightness law that cannot be gamed at all still hurts, most likely because straightening an early, blurry estimate pushes the model toward futures in which nothing ever collides.

The same laws are useful in a different role. Instead of guiding, they score finished clips, and the best-scoring one is kept. Nothing is being optimised against them, so there is nothing to game, and they rank clips well.

The renoise-and-denoise rounds help every method except block 4

Download PNG
Local error against clip length for six methods, each drawn twice: with 24 renoise-and-denoise rounds and with none. Every method except block 4 is lower with the rounds; block 4 is lower without them. The full law set is flat across clip length with the rounds and rises without them
Figure 46. Each method appears twice, once with 24 renoise-and-denoise rounds and once with none, and nothing else differs between the pair. A round mixes the finished clip back to σ 0.45 and runs the last 22 of the 50 denoising steps again. Solid is whichever setting suits that method. Every curve is the mean error of 8 independent samples: the sampler runs 8 times from 8 different starting noises and the 8 error numbers are averaged, so no clip is selected and no trajectory is averaged with another. The guided curves run the model twice per denoising step, at a strength picked separately for each setting on the 128 calibration clips. Hollow markers lost more than 5% of the data’s mean speed, so that error is partly bought by slowing the balls down instead of predicting better.
Δx(5)rounds25f33f41f49f25→49keeps the data’s speed?
oracle (positive control)240.0980.0970.0960.0890.91×every length
none0.1110.1100.1090.1020.92×every length
block 4 (autoregressive)240.2660.3040.3200.3271.23×no, at 49 frames
none0.2410.2410.2360.2250.93×every length
full law set240.3340.3320.3330.3220.97×every length
none0.4800.5350.5500.5421.13×no, at 49 frames
rollout-5 target240.3460.4140.4780.4981.44×no, from 41 frames
none0.4350.5430.6060.6241.43×no, at any length
rollout-1 target240.4090.4890.5300.5371.31×no, from 33 frames
none0.5420.6490.6890.6951.28×no, at any length
vanilla240.5720.6690.6890.6911.21×every length
none0.6880.8340.8820.8911.29×every length

896 holdout clips, mean over 3 training seeds and over 8 independent samples per clip. Speed is the mean ball speed measured from the generated positions, divided by the truth’s; a method that slows every ball down lowers the local error without predicting better, so a value below 0.95 disqualifies the number.

The rounds are what makes the law set competitive, and what flattens its curve

Take the rounds away and the full law set loses to block 4 by a wide margin at every clip length. It also stops being flat: with the rounds its error is the same on the longest clips as on the shortest, and without them it grows like everything else. The flat curve is therefore a property of the laws and the rounds together, not of the laws alone. This is the sharpest statement the section can make about the law set, and it is weaker than the one the earlier comparison suggested.

Block 4 is the only method the rounds damage

Every other method is more accurate with them, the unguided model most of all. Block 4 gets worse at every clip length, and on the longest clips it also stops keeping the data’s speed. A round re-noises the whole clip and regenerates it block by block, which breaks the agreement between neighbouring blocks that makes generating in order work. That is why block 4 is drawn solid without the rounds: that is its best setting, and it is the number the guided methods have to beat.

The rounds do not rescue the rollout targets

Both rollout targets are already slowing the balls down without the rounds, at every clip length, so none of their numbers count in either setting. Adding the rounds lowers their error but does not fix the speed, and on the longer clips it is still failing. The oracle, which is given the true future, keeps the data’s speed throughout and is unaffected either way.

The same comparison with every method keeping the best of its eight independent samples

Download PNG
The previous comparison with one setting changed: every method now keeps the clip its law critic scores best out of eight independent samples. Every curve drops, the ordering barely moves, and block 4 given the renoise-and-denoise rounds no longer carries a hollow marker
Figure 47. Every method here gets the 24 renoise-and-denoise rounds, and one setting is added on top: instead of averaging over its eight independent samples, each method keeps the one its law critic scores best. The critic scores a finished clip on three of the laws — position-only free flight, speed constancy outside contacts, and total momentum outside wall contacts — fixed in advance on a separate run. The dashed lines are the same methods averaging over their eight samples instead, which is what Figure 46 plots. The distance from a solid line to its dashed partner is what the selection buys, and it is similar for every method.
Δx(5), 24 renoise-and-denoise rounds, best of 8 independent samples25f33f41f49f25→49keeps the data’s speed?
oracle (positive control)0.0960.0960.0950.0880.91×every length
block 4 (autoregressive)0.2370.2610.2670.2761.17×every length
full law set — the best we found0.2930.2850.2860.2780.95×every length
geometry laws only (superseded)0.3650.3870.3910.3881.06×every length
rollout-5 target0.3070.3550.4060.4251.38×no, from 41 frames
rollout-1 target0.3500.4120.4580.4741.35×no, from 33 frames
vanilla0.5560.6380.6530.6611.19×every length
block 4, no renoise-and-denoise — its best setting0.2410.2410.2360.2250.96×every length

896 holdout clips, mean over 3 seeds.

Keeping the best of eight helps every method, and helps them by similar amounts, so the ordering barely changes. Whatever the critic is picking up on is not specific to the law-guided model. The full law set and block 4 given the same renoise-and-denoise rounds finish level on the longest clips either way, and block 4 run plainly stays clearly ahead of both.

One thing does change. Block 4 given the renoise-and-denoise rounds fails the speed check at the longest clip length when its eight independent samples are averaged, and passes it when the critic chooses one. The selection is partly repairing what the renoise-and-denoise rounds do to that model — more evidence that the renoise-and-denoise rounds are the wrong setting for it, and that its plain run is the number worth comparing against.

What the full law set actually does

At every denoising step the model produces its current guess at the finished clip. Ten laws are evaluated on that guess, each one an equation the true motion satisfies exactly, and the guess is nudged toward satisfying them. Nothing simulates the motion forward at any point.

Write \(p_t^b\) and \(v_t^b\) for the position and velocity of ball \(b\) at frame \(t\) in the current guess, \(\Delta t\) for the frame interval, \(r\) for the ball radius, and \(\rho_t^{ij} = p_t^i - p_t^j\) for the vector between a pair of balls. Every law is written on one interval, from frame \(t-1\) to frame \(t\), and averaged over balls or pairs and over the generated frames.

Which balls each law applies to

A law about free flight must not be applied to a ball that just collided, and a law about collisions must not be applied to a ball that did not. Neither is known, so both are estimated from the guess itself. For a pair, three distances are computed: the closest approach if both balls fly freely forward from \(t-1\), the closest approach flying backward from \(t\), and the closest approach of the straight segment joining the two frames. Taking the smallest, \(d^{ij}\), a soft indicator of contact is

\[ s^{ij} \;=\; \sigma\!\left(\frac{2r + m - d^{ij}}{\tau}\right), \]

and the chance that ball \(b\) touched nothing in the interval is \(\phi^b = \prod_{j \neq b}\left(1 - s^{bj}\right)\). The free-flight laws are weighted by \(\phi^b\), the collision laws by \(s^{ij}\). The margin \(m\) is set wider for the free-flight laws than for the collision laws, so that a collision the indicator underestimates does not leave a free-flight law fighting it.

The ten laws

Geometry, which holds at every frame regardless of what happened:

\[\begin{aligned} L_{\text{overlap}} &= \big[\,2r - \lVert \rho_t^{ij} \rVert\,\big]_+^2 &&\text{no two balls overlap}\\ L_{\text{through}} &= \big[\,2r - d^{ij}_{\text{seg}}\,\big]_+^2 &&\text{no ball crosses through another}\\ L_{\text{box}} &= \textstyle\sum_{\text{axes}} \big[\,\ell - p_t^b\,\big]_+^2 + \big[\,p_t^b - u\,\big]_+^2 &&\text{no ball leaves the box} \end{aligned}\]

Free flight, for a ball the indicator says touched nothing. Let \(q = p_{t-1}^b + v_{t-1}^b \Delta t\), mirrored once about a wall if it lands outside, and \(\tilde v = -v_{t-1}^b\) if it was mirrored and \(v_{t-1}^b\) otherwise:

\[\begin{aligned} L_{\text{pos}} &= \phi^b \, \lVert p_t^b - q \rVert^2 &&\text{it lands where a straight line puts it}\\ L_{\text{vel}} &= \phi^b \, \lVert v_t^b - \tilde v \rVert^2 &&\text{its velocity is unchanged, or reflected by the wall} \end{aligned}\]

Collisions, for a pair the indicator says touched. Write \(V = v^i + v^j\) for the pair’s total velocity and \(E = \lVert v^i\rVert^2 + \lVert v^j\rVert^2\) for its energy:

\[\begin{aligned} L_{\text{mom}} &= s^{ij}\,\lVert V_t - V_{t-1} \rVert^2 &&\text{momentum is conserved}\\ L_{\text{energy}} &= s^{ij}\left(\frac{E_t - E_{t-1}}{E_{t-1}}\right)^{\!2} &&\text{energy is conserved}\\ L_{\text{com}} &= s^{ij}\,\big\lVert (p_t^i + p_t^j) - (p_{t-1}^i + p_{t-1}^j) - V_{t-1}\Delta t \big\rVert^2 &&\text{the centre of mass keeps its velocity} \end{aligned}\]

The last two need the contact normal. Flying both balls freely forward from \(t-1\), let \(s_c \in (0,1)\) be the fraction of the interval at which their separation first reaches \(2r\), \(\rho_c\) the relative position there, and \(n = \rho_c / \lVert \rho_c \rVert\). For equal masses an elastic collision reflects the relative velocity about \(n\), giving \(\hat u = u_{t-1} - 2(u_{t-1}\!\cdot n)\,n\) where \(u = v^i - v^j\):

\[\begin{aligned} L_{\text{reflect}} &= s^{ij}\,\lVert u_t - \hat u \rVert^2 &&\text{the relative velocity reflects about the contact normal}\\ L_{\text{cpos}} &= s^{ij}\,\big\lVert \rho_t^{ij} - \big(\rho_c + \hat u\,\Delta t\,(1 - s_c)\big) \big\rVert^2 &&\text{and the pair ends the interval where that puts it} \end{aligned}\]
Dividing by the current speed, which is what made this set work

As written, every residual above shrinks if the balls simply move less. A guidance that slows everything down therefore lowers the loss without predicting anything better, and that is what an earlier law set did. So each residual is divided by how fast the balls are currently moving — \(\lVert v_{t-1}^b \rVert + \varepsilon\) for a single ball, the sum of both speeds for a pair — squared to match the residual. Scaling every velocity in the clip by a constant now leaves the loss unchanged, so slowing down buys nothing.

How the laws become a correction

The laws are not summed into one loss and differentiated once. Each law’s gradient with respect to the current guess is taken separately, and each gets its own step, because the gradients differ in scale by orders of magnitude: a law that is already satisfied almost everywhere touches a handful of numbers per frame, and adding it to a dense one would bury it. For a law taking a bounded step, the gradient is divided by its own largest element and scaled by that law’s weight, so no law can move any single number by more than its share. For the rest, the step is the exact minimiser along the gradient, which for a squared residual is

\[ \alpha_k = w_k\,\frac{L_k}{\lVert \nabla L_k \rVert^2}, \qquad \text{correction} = \alpha_k \nabla L_k, \]

capped so one law cannot dominate a step. The per-law corrections are then added. This happens twice per denoising step.

Two things sit on top and are not guidance: the finished clip is partly re-noised and denoised again twenty-four times, and several clips are generated with three of the laws used to score them and the best kept. Both help any method, which is what the comparison above controls for.

Most of the apparent win came from the extra sampling work

The geometry law set looked far better than the rollout target, but the rollout target had been run without renoise-and-denoise or repeated independent samples. Give it both and almost all of that difference disappears. Those two settings help any guidance, and they help the unguided model most of all. The full law set does open a genuine gap over the rollout target, and that gap comes from the laws, because both methods now get the same sampling.

The full law set beats both rollout targets at every clip length

Given the renoise-and-denoise rounds it does so whether each method averages its eight independent samples or keeps the best of them, and the margin widens as clips get longer. The rollout targets also slow the balls down, so their numbers stop counting from the middle clip lengths on, and without the rounds they fail that check at every length. The law set keeps the data’s speed everywhere except the longest clips without the rounds.

Law guidance stops the error growing with clip length, but only with the renoise-and-denoise rounds

Given the rounds, the law set’s error is the same on the longest clips as on the shortest, and every other guided method gets worse as clips lengthen. Take the rounds away and the law set grows with length like the rest. So the flat curve needs both the laws and the rounds; neither produces it alone. Selection is not involved either way: the curve is flat with a single sample and with the best of eight.

Given the same sampling, law guidance matches generating in order

On the longest clips the best law set and block 4 given the same renoise-and-denoise rounds reach the same five-step error, within the spread across training seeds, and the law set is ahead on the one-step error. Without the selection, block 4’s number at that length is disqualified by the speed check as well. The two cross: block 4 leads on short clips and loses that lead on long ones, because the renoise-and-denoise rounds make block 4 worse as clips lengthen while the law set holds steady.

That tie depends on holding block 4 to settings that damage it. A renoise-and-denoise round mixes a finished clip back into noise and regenerates it block by block, which breaks the agreement between neighbouring blocks that makes block 4 work. Run without renoise-and-denoise it is ahead on both errors, and the law-guided model cannot match it there because it needs the renoise-and-denoise rounds to reach its own best result. So: equalise the sampling and the two tie; let each method use the settings that suit it and generating in order still wins.

The first law set’s improving curve came from slower balls

The first law set, the one with the differentiable simulator, appeared to get better on longer clips. Its balls also move progressively slower than real ones as clips lengthen. The error measures how far a ball has drifted from where it should be, and a slower ball drifts less, so the error falls without the prediction improving. That curve is not evidence of anything and is no longer plotted.

Where the remaining error is: one generated clip

Download PNG
A grid of 44 frames by 5 balls for one law-guided clip. Cell colour is the five-step error at that frame and ball, white to dark red. Hatched cells are windows in which at least one frame interval has no applicable law. Nearly every dark cell is hatched and nearly every white cell is not.
Figure 48. One clip generated with the position-only laws (2 renoise-and-denoise rounds, then 10 iterations of the laws on the finished clip). Each cell is one ball at one frame t; its colour is that ball’s five-step error at t, which is computed from the five frame intervals t−5 … t. A cell is hatched when at least one of those five intervals is one where no law applies to that ball. Unhatched cells: mean error 0.06. Hatched cells: 0.57.

The same, over 256 clips

Download PNG
Three panels. (a) For the position-only laws, 12% of windows are covered and they hold 2% of the five-step error; for the velocity-channel laws, 38% of windows are covered and hold 19% of the error. (b) The uncovered windows split by what the simulator finds in them: no contact 23% of the error, a wall bounce only 16%, one collision 34%, two or more collisions 24%. (c) Ten iterations of the position laws on the finished sample divide their own loss by 18 while the five-step error goes from 0.525 to 0.517.
Figure 49. 49 frames, seed 0, 256 held-out clips. (a) Left bar: the share of (frame, ball) windows in which a law applies in all five intervals. Right bar: the share of the total five-step error that sits in those windows (green) and in the rest (red). (b) The uncovered windows, split by what the simulator finds when it runs the five frames from the generated state. (c) Iterating the position-only laws on the finished clip: their own loss falls 18-fold; Δx(5) moves by 0.008.

The laws are satisfied where they apply; the error is where they do not apply

Every law comes with a condition for when it may be used. Free flight applies to a ball only if no other ball and no wall is within one frame of travel (0.6 units) of its path, because otherwise a contact may fall inside the interval and a straight line would be the wrong prediction. The collision laws apply to a pair only if the two balls touch and nothing else is within reach of either, because the outcome of a contact next to a wall or a third ball depends on which contact comes first. Where the condition fails, the law says nothing.

Figure 48 shows what that leaves. For the exact position-only laws, 12% of the (frame, ball) windows have a law applying in all five of their intervals. Those windows are nearly right already (0.06 against 0.02 for the true clips) and hold 2% of the total error. The other 98% of the error is in windows with at least one interval the laws leave alone (panel a). Panel b says what is in those intervals: a collision with something else within reach (34% of the error), two or more collisions in the five frames (24%), a wall bounce with another ball within reach (16%), and near passes where nothing touches but a ball or wall was within one frame of travel, so no law was sure enough to apply (23%).

Panel c is the direct test. Ten more iterations of the position laws on the finished clip lower the laws’ own loss 18-fold and the five-step error by 0.008. The laws are being satisfied; the score barely notices, because the score lives in the intervals the laws do not touch.

The velocity-channel laws have looser conditions and cover 38% of the windows, but the intervals they leave hold 81% of the error, and even in the windows they cover the guided clip is 1.7× worse than block 4 (0.17 against 0.10): those laws constrain the generated velocities, and the score reads only positions, so a clip can satisfy them by adjusting its velocity numbers alone. Block 4 has the same 10% coverage and the same split (3% of its error in covered windows), and is better in both kinds of window, 0.07 and 0.24 against 0.10 and 0.57.

What is left uncovered has one thing in common: two events close together, and the outcome depends on their order. A law for that case must compute which of two contacts happens first from the state at the start of the interval, apply it, and then check the second. That is the simulator’s own step. The law sets above stop at 0.32 because every law in them assumes one event per interval, and the error that remains is in the intervals that break that assumption.

What the local errors do not show

The plausibility measures from earlier, recomputed on the trajectories both law sets produced:

49 framestruthvanillarollout-5oracleblock 4laws + rolloutslaws only
Δx(5)00.8880.6390.1880.2230.3170.388
frames with overlapping balls00.1400.0710.0040.0080.0000.000
collision-law residual, ball–ball00.7060.5760.1300.2290.2930.335
phantom velocity changes00.0670.1360.0170.0140.0820.035
final energy1.000.950.470.980.980.540.95
ball–ball collisions per clip15.125.015.415.114.75.511.1
W1 distance, speeds00.231.720.060.111.660.21
median frame of divergence49131549181313

The first law set gets part of its accuracy by avoiding collisions rather than by getting them right. Its clips contain roughly a third as many ball-to-ball collisions as real ones, and the balls lose half their energy. A collision that never happens cannot be resolved wrongly, so dodging them lowers the error without improving the physics.

The geometry law set does not cheat that way. It keeps as much energy as the unguided model, matches the real distribution of ball speeds, halves the unguided model’s unexplained velocity changes, never overlaps two balls, and halves its collision-law error. It still collides less often than real clips, which is its one remaining defect — milder than the first law set’s, and absent from the rollout target, which gets the collision rate right.

Neither law set improves the global error. Both produce a physically consistent future rather than the true one. Block 4 is still the only method that is accurate, keeps its energy and collides as often as real clips, all at the same time.

Cost and data

What was run

882 runs for the main grid, ablations and strength sweep on one node of RTX PRO 6000 GPUs, 48 at a time: about 40 minutes wall, 10–20 s per unguided or oracle run and 40–150 s per rollout run, where the simulator in a 16-process pool dominates. The physics-law search added 239 runs over eight iterations, 20 s for single laws to 17 min for the final configuration with 16 renoise-and-denoise rounds, plus about 40 CPU screens on 64 clips. The whole-trajectory distribution comparison is 72 cells, about 35 minutes on 40 CPUs and no GPU.

Per-run and per-clip data

The selected strengths and aggregated metrics, every individual run, and the two evaluation suites.

Off the path · Not run yet

To-do & ideas

CollapseExpand

Things worth running that are not part of the sequence above. Nothing here has been measured; each entry says what it would show and where it would hook into the existing code.

Idea 1 · Visualization

Watch each frame settle over the denoising steps

Show frames 1, 2, 4, 10, 20, 30 and 40 of one generated clip side by side — but animate each panel over the denoising steps instead of over clip time. Panel t plays \(k=0,\dots,K\): the state frame t holds at each step of the reverse process. The clip’s own motion is frozen; what moves is the model’s guess about that one frame.

What it would show. Where in the loop each frame stops changing. If the late frames keep moving until the last few steps while the early ones settle immediately, the loop is spending its steps unevenly along the clip. The quantitative version of the same picture is per-frame displacement per step, \(\lVert \hat x_k(t)-\hat x_{k-1}(t)\rVert\) against \(k\), one curve per frame.

Where it hooks in: sample() in experiments/statebench/statebench/diffusion.py already holds the running state at every one of the K steps. Record the clean estimate \(\hat x = x_\sigma - \sigma v\) at each step rather than only the final state, then render the selected frames with the existing trail renderer, one video per frame with the denoising step as the time axis. Worth doing for both models: under block 4 a frame should sit still except during its own block, which is the same contrast section 03 reports as numbers.

Idea 2 · Measurement

Look at the score direction from a cloud of neighbouring points

Take one noisy state \(x_k\) at a fixed noise level and evaluate the model not only there but at many neighbours \(x_k+\delta\), with \(\delta\) drawn from a small sphere. Then compare the directions the model returns.

What it would show. Whether the model’s error is coherent. With an exact score the field around \(x_k\) is affine: \(s^*(z)=\big((1-\sigma)y^*-z\big)/\sigma^2\), so every neighbour implies the same clean future \(y^*\), and the Jacobian is exactly \(-I/\sigma^2\). Two things are then measurable on a cloud of neighbours — the spread of the implied clean states \(\hat x(z)\), and the finite-difference Jacobian against \(-I/\sigma^2\). If the neighbours all point at one wrong future, the model has learned a coherent but incorrect target; if they disagree, there is no single target at all. Section 01’s argument is about that Jacobian, and section 03.5 measures only the size of the score error, never its local structure.

Cheap to run: one batch of perturbed copies of a single clip through the network, at a few noise levels. The exact score for every neighbour is available in closed form, because the conditioning frames fix \(y^*\).