machine learning

The Seriality Gap in Video Diffusion Models

arXiv:2607.13031

summary

The paper investigates why video diffusion models struggle with tasks that require sequential causal reasoning, such as multi‑ball collisions, and identifies a "seriality gap" where increasing denoising steps does not provide the needed serial computation.

Abstract

When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched single-ball control, where ball-ball interactions are absent, the degradation largely disappears, isolating dependent-event structure rather than video length as the cause. Across intervention studies, methods that increase effective serial computation improve performance disproportionately, including autoregressive/blockwise generation and architectural depth. We identify this pattern as the seriality gap: a mismatch between tasks requiring growing serial computation and video diffusion models whose denoising loop does not provide scalable serial compute. We then prove that, for deterministic video prediction, denoising steps do not add serial computation beyond the backbone, indicating a structural obstacle for video diffusion on serial reasoning and simulation tasks.

Jorge Diaz Chao and Konpat Preechakul contributed equally. 24 pages, 12 figures, and 5 tables. Project page: https://seriality-gap.jdiazchao.com

Topics & keywords

#video diffusion#causal reasoning#serial computation#simulation#autoregressive generationdenoising diffusion probabilistic modelsbidirectional diffusionautoregressive blockwise generationdeterministic video predictionseriality gap
The Seriality Gap in Video Diffusion Models · wovepaper