5 papers
EventFlash: Towards Efficient MLLMs for Event-Based Vision
Shaoyu Liu, Jianing Li, Guanghui Zhao +4
Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, curre…
LoL: Longer than Longer, Scaling Video Generation to Hour
Justin Cui, Jie Wu, Ming Li +6
Recent research in long-form video generation has shifted from bidirectional to autoregressive models, yet these methods commonly suffer from error accumulation and a loss of long-…
Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
Justin Cui, Jie Wu, Ming Li +6
Diffusion models have revolutionized image and video generation, achieving unprecedented visual quality. However, their reliance on transformer architectures incurs prohibitively h…
RewardDance: Reward Scaling in Visual Generation
Jie Wu, Yu Gao, Zilyu Ye +9
Reward Models (RMs) are critical for improving generation models via Reinforcement Learning (RL), yet the RM scaling paradigm in visual generation remains largely unexplored. It pr…
Efficient Video to Audio Mapper with Visual Scene Detection
Mingjing Yi, Ming Li
Video-to-audio (V2A) generation aims to produce corresponding audio given silent video inputs. This task is particularly challenging due to the cross-modality and sequential nature…