#video diffusion
5 papers match
Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion
Henglin Liu, Fangyuan Kong, Jing Wang +7
The paper introduces concentrated Implicit Preference Optimization (cIPO), a post‑training method for text‑to‑video diffusion models that derives preference signals from reconstruc…
AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight
Xinhong Zhang, Qiyuan Zhu, Yubo Huang +8
The paper introduces AeroAct, a world-action model that predicts quadrotor flight actions from egocentric video, proprioceptive data, and language commands, using a video diffusion…
Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data
Thang-Anh-Quan Nguyen, Moussab Bennehar, Luis Guillermo Roldao Jimenez +5
Cyclone is a latent diffusion framework that edits weather conditions in driving images without paired data, using cycle-consistent constraints and image‑text knowledge to produce…
WanToFight: Real-Time Generative Game Engine for Multi-Player Combat Interaction
Li Hu, Guangyuan Wang, Peng Zhang +1
WanToFight is a generative game engine that uses a video diffusion transformer to produce real-time, two-player fighting game visuals from keyboard inputs, handling multi-player co…
The Seriality Gap in Video Diffusion Models
Jorge Diaz Chao, Konpat Preechakul, Yuxi Liu +1
The paper investigates why video diffusion models struggle with tasks that require sequential causal reasoning, such as multi‑ball collisions, and identifies a "seriality gap" wher…