← Back to rollout-5 guidance results

Full experiment report

Rollout-5 guidance experiments

This report reconstructs the rollout-5 guidance work in out/guidance and evaluates saved outputs with the original rollout-5 metric from this repository. Ball validity is shown separately. No figure uses the later penalized rollout-5 error.

The experiments progressed from broad latent-guidance tests to increasingly direct and selective corrections. The broad tests established that the sampler responds strongly to an exact future target, but rollout targets built from the generated video did not improve physical consistency. Stronger corrections made this worse. Later work removed ground-truth inputs, built continuous physics windows, checked proposals before accepting them, waited for clean late-step detections, and changed the latent target. Those changes often restored missing or incorrect balls, but all six five-video variants raised mean original rollout-5 error. A final 60-run sweep on one Case 2 generation produced one saved output with a 7.2% lower original rollout-5 mean than its control. That result was not tested across more videos.

1. Metric and scope

The original metric reconstructs the generated trajectory with Simulation.from_video. For each target frame from 5 through 48, it starts five frames earlier, simulates forward for five frames, and compares the simulated and reconstructed ball centers. The reported value is the mean Euclidean distance over 44 target frames × 5 balls = 220 comparisons, in the simulator's 10-by-10 world coordinates.

The exact calculation is:

Difference.from_simulations(
    simulation,
    rollout(simulation, horizon=5),
).avg_center_error((5, None))

See the evaluation code and rollout implementation.

The separate scorable-pair count comes from the stricter tracker introduced during these experiments. Complete coverage is 220 pairs per video. Missing balls, wrong-color balls, extra balls, or an unavailable source state reduce that count. It describes ball validity and does not alter the original rollout-5 value.

Two data sources are used below:

  • The initial holdout summaries already contain the original rollout_h5_error produced by the repository evaluator.
  • Later MP4s were rescored directly. The new score archive covers 136 saved videos. Thirty-two identical files overlap the independently validated rollout-1 report archive; all 32 rollout-5 values match exactly.

2. The initial matrix isolated the target as the main problem

The first large experiment compared ordinary diffusion with four interventions across six model conditions, five reverse-step counts, and three seeds. Each seed evaluated 896 videos.

Method Correction
Gaussian control Add a random latent correction of matched size.
Absolute rollout target Decode the Tweedie estimate, construct a five-frame physics rollout using the exact evaluation initial state, encode the complete reference, and pull toward it.
Physics-delta target Use only the latent displacement caused by the physics rollout, removing the encode/decode round-trip offset.
Exact-future oracle Pull toward the encoded future from the target video.

The table gives percent improvement in the original rollout-5 error, averaged over 50, 100, and 200 reverse steps. Positive values mean lower error.

Condition Gaussian Absolute rollout Physics delta Exact-future oracle
1 ball, 25 frames -0.02% -0.40% +0.09% +36.68%
1 ball, 49 frames -0.01% +2.07% -0.01% +39.74%
5 balls, 25 frames -0.24% +0.72% -0.65% +16.12%
5 balls, 33 frames -0.10% +0.16% -0.10% +19.38%
5 balls, 41 frames +0.03% -0.14% -0.15% +32.23%
5 balls, 49 frames +0.16% +0.08% +0.12% +34.92%

Initial holdout matrix using original rollout-5

The exact-future oracle improved all six conditions by 16–40%. Gaussian noise and both rollout-derived targets stayed within 2.1% of diffusion and changed sign across conditions. The sampler could use a good direction. The rollout construction did not supply one consistently.

Sources: absolute-target holdout, physics-delta holdout, and combined report data.

3. Increasing the correction strength exposed a bad direction

The calibrated physics-delta strength, 0.003, barely changed the videos. A four-video, five-ball, 49-frame sweep increased both absolute-target and delta-target strengths from 0.003 to 0.3.

Strength sweep using original rollout-5

The four diffusion controls average 0.9843 original rollout-5 error. The lowest tested absolute-target mean is 1.0218 at strength 0.01. The lowest physics-delta mean is 0.9975 at strength 0.003. At strength 0.3, the means rise to 1.7015 and 2.1062.

Larger updates made the effect visible but did not repair it. This motivated changes to the reference and state extraction rather than another scale increase.

Source: rescored strength sweep.

4. Cleaner state handling and local trajectories reduced known defects, but not rollout-5

The next implementation removed evaluation-simulator metadata. It detected balls in the current Tweedie estimate, used the state exactly five frames before each target, and built a complete reference from separately simulated endpoints. This made the guide self-contained, but its mean original rollout-5 error rose from 1.0347 to 1.1666 on four saved matched videos.

The following version replaced the stitched endpoints with one continuous five-frame trajectory per source and applied the latent change only to that local window. On four saved matched videos, the mean rose from 0.9216 to 0.9292.

The archived complete run for this version contains 1,024 paired generations, but its summary stores only the later tracker metrics. Its original rollout-5 values cannot be reconstructed because those full-run videos are absent. The four surviving matched pairs are the original-metric evidence used here.

The next step tried several strengths at the same noisy latent and accepted only a proposal that lowered the experiment's guidance score. It was tested twice on Case 1. The original rollout-5 results were:

Run Vanilla Guided Change Accepted updates
Seed 0 1.0273 1.0743 +0.0470 33
Seed 1 1.0607 1.0132 -0.0475 36
Mean 1.0440 1.0437 -0.0003 69 total

Successive reference and acceptance changes using original rollout-5

The three stages use different saved video sets, so the useful comparison is the guided-minus-control change within each stage. Cleaning the reference made three of four videos worse. Continuous trajectories made three of four worse by smaller amounts. Score checking split evenly across the two Case 1 generations and left their mean unchanged.

Sources: paired construction results, Case 1 run metadata (out/guidance/continuous-window-accept-case1-v1), and historical implementation report (out/guidance/guidance-experiments-report/index.html).

5. Late guidance tested every usable five-frame window

Early Tweedie estimates often lacked clean detections. The next set waited until reverse steps 40–49, found every usable source frame, simulated each five-frame future, combined the latent losses, and allowed up to 12 accepted updates per denoising step. Every proposal tested strengths 0.1, 0.03, 0.01, and 0.003.

This stage changed one construction choice at a time:

  1. Apply the physics-delta loss to the full five-frame latent window.
  2. Apply it only to the latent time entry containing the endpoint.
  3. Reject a source frame when the detector finds an extra ball.
  4. Replace the delta target with the directly encoded physics target.

All six variants were run on the same five source videos and generation seed. Their common vanilla mean was 1.4163.

Variant Guided mean Change Videos improved Accepted updates Scorable pairs, before → after
Delta, full window 1.4850 +0.0687 2/5 91 890 → 903
Delta, endpoint 1.5192 +0.1028 2/5 85 890 → 897
Delta, full + strict validity 1.4540 +0.0377 1/5 80 843 → 856
Delta, endpoint + strict validity 1.4745 +0.0581 2/5 98 843 → 910
Direct target, full window 1.4579 +0.0415 1/5 93 843 → 1,033
Direct target, endpoint 1.5006 +0.0843 2/5 97 843 → 961

Per-case changes for late guidance variants using original rollout-5

Every variant increased the five-video mean. Case 4 became worse under every method. Case 5 improved under five methods, including a 0.2222 reduction from the direct full-window target. This variation did not identify a setting that transferred across the five videos.

The direct target produced the largest validity gain. Full-window guidance added 190 scorable pairs, reaching 1,033/1,100, while increasing original rollout-5 by 0.0415. The other variants show the same separation between validity and the original metric.

Validity change against original rollout-5 change

This explains why the historical acceptance score could improve while the original rollout-5 mean worsened. The two measurements reward different behavior: one strongly tracks whether balls remain valid, while the original metric measures the simulator mismatch after the repository tracker reconstructs the trajectory.

Sources: per-case original scores and validity, late-window runs (out/guidance/late-all-window-local-5cases), endpoint runs (out/guidance/late-endpoint-local-5cases), strict-validity runs (out/guidance/late-extra-invalid-full-window-5cases), and direct-target runs (out/guidance/late-absolute-full-window-5cases).

6. The focused Case 2 sweep found a single original-metric improvement

The final rollout-5 investigation changed the update rule again. It repeatedly selected the smallest source error above a threshold and applied direct gradients without proposal rollback. Sixty full horizon-5 generations tested strength, start step, threshold, gradient scope, update count, source order, and correction normalization.

This entire sweep used Case 2, 00000/828.mp4, generation seed 0, five observed frames, 49 output frames, and 50 denoising steps. Only two guided horizon-5 videos survive in the report archive, so only those two can be evaluated with the original metric.

Saved output Key configuration Original rollout-5 p95 Maximum Scorable / 220
Vanilla No guidance 1.4111 4.0851 7.6662 55
Visual-balance selection Start 25; 50 updates/step; strength 0.20 → 0.06 1.3099 3.0332 6.2728 204
Historical score selection Start 25; 75 updates/step; strength 0.225 → 0.0675; ordered sources 1.4508 4.2605 7.3904 220

Focused Case 2 selected outputs using original rollout-5

The visual-balance output lowers the original mean by 7.2%, p95 by 25.7%, and maximum by 18.2% relative to vanilla. It also raises validity from 55 to 204 scorable pairs.

The configuration selected historically for its later score gives complete validity, but its original rollout-5 mean is 2.8% higher than vanilla and its p95 is 4.3% higher. The selection order therefore changes when the original metric is used.

Decoded frames from the saved Case 2 outputs

The 58 other horizon-5 final videos are absent. Their saved summaries use the later metric and cannot determine their original rollout-5 ranking. The measured 7.2% gain belongs to one generation; no multi-video run tested this configuration.

Sources: rescored selected outputs, historical Case 2 report (out/guidance/guidance-experiments-report/rollout-smallest-error-results-so-far.html), visual-balance video, and historical score-selection video.

7. Result

The broad result is negative for rollout-derived guidance under the original rollout-5 metric. The initial holdout matrix found no consistent gain, the strength sweep worsened as corrections grew, and all six late five-video variants increased the mean. The exact-future oracle improved every model condition, so the sampler itself can respond to a useful latent direction.

The focused Case 2 sweep supplies the one positive final-video result: the saved visual-balance output improves original rollout-5 by 7.2% and greatly improves validity. It is evidence that direct rollout-5 guidance can help a particular generation. It is not evidence of improvement across videos because the configuration was only evaluated on Case 2.

Data and reproduction

No new videos were generated for this report.

python out/guidance/rollout5-report-assets/score_original_rollouts.py
python out/guidance/rollout5-report-assets/build_figures.py

Return to the results →