7 papers · 1 filter
Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention
Zekun Li, Xiaoyan Cong, Hongyu Li +5
Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional n…
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch +2
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements…
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
Xiaoyan Cong, Zekun Li, Zhiyang Dou +9
Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets…
PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
Aashish Rai, Angela Xing, Anushka Agarwal +5
Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality re…
HFGaussian: Learning Generalizable Gaussian Human with Integrated Human Features
Arnab Dey, Cheng-You Lu, Andrew I. Comport +3
Recent advancements in radiance field rendering show promising results in 3D scene representation, where Gaussian splatting-based techniques emerge as state-of-the-art due to their…
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
Aashish Rai, Srinath Sridhar
We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric v…