The Seriality Gap in Video Diffusion Models
Jorge Diaz Chao, Konpat Preechakul, Yuxi Liu, Yutong Bai
Read on arXiv →Key claim
Video diffusion models face a seriality gap in performance.
In plain English
Video models struggle to predict outcomes in scenarios with multiple interacting objects, especially as the number of interactions increases. Current video diffusion methods fail to scale effectively with the complexity of these tasks, leading to performance degradation. This paper highlights the importance of serial computation in improving model performance and suggests methods to enhance it. Builders should care because addressing this seriality gap could lead to more robust video prediction systems that better simulate real-world dynamics.
Introduces the concept of the seriality gap in video diffusion models.
Presents controlled experiments and intervention studies to support claims.
Deep reliability assessment
The methodology supports the claim that video diffusion models struggle with serial computation in deterministic settings, but it may overclaim the generality of this limitation across all video prediction tasks.
Reproducibility
No open source code or dataset is mentioned in the paper.
Key figure
Figure 1 illustrates the difference between non-serial and serial video prediction using hard-sphere dynamics, highlighting the challenge of dependent-event chains.
