6 papers · 1 filter
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
Ruicheng Zhang, Kaixi Cong, Jun Zhou +5
Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration…
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
Zishan Liu, Ruoxi Zang, Yanglin Zhang +5
Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, ofte…
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
Zhiwei Ning, Xuanang Gao, Jiaxi Cao +6
Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a persistent challenge. Although re…
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
Yicheng Ji, Zhizhou Zhong, Jun Zhang +7
Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness, as exemplified by the Self…
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
Haozhe Liu, Ding Liu, Mingchen Zhuge +16
We introduce MoS (Mixture of States), a novel fusion paradigm for multimodal diffusion models that merges modalities using flexible, state-based interactions. The core of MoS is a…
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero +81
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent…