8 papers
Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention
Zekun Li, Xiaoyan Cong, Hongyu Li +5
Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional n…
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Rahul Sajnani, Yulia Gryaditskaya, RadomÃr MÄch +2
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements…
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
Xiaoyan Cong, Zekun Li, Zhiyang Dou +9
Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets…
PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
Aashish Rai, Angela Xing, Anushka Agarwal +5
Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality re…
MotionGlot: A Multi-Embodied Motion Generation Model
Sudarshan Harithas, Srinath Sridhar
This paper introduces MotionGlot, a model that can generate motion across multiple embodiments with different action dimensions, such as quadruped robots and human bodies. By lever…
GeoDiffuser: Geometry-Based Image Editing with Diffusion Models
Rahul Sajnani, Jeroen Vanbaar, Jie Min +2
The success of image generative models has enabled us to build methods that can edit images based on text or other user input. However, these methods are bespoke, imprecise, requir…