collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

Zekun Li, Xiaoyan Cong, Hongyu Li +5

Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional n…

cs.CV2026

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch +2

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements…

cs.CV2026

UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors

Xiaoyan Cong, Zekun Li, Zhiyang Dou +9

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets…

cs.CV2026

PackUV: Packed Gaussian UV Maps for 4D Volumetric Video

Aashish Rai, Angela Xing, Anushka Agarwal +5

Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality re…

cs.CV2024

HFGaussian: Learning Generalizable Gaussian Human with Integrated Human Features

Arnab Dey, Cheng-You Lu, Andrew I. Comport +3

Recent advancements in radiance field rendering show promising results in 3D scene representation, where Gaussian splatting-based techniques emerge as state-of-the-art due to their…

cs.CV2024

EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos

Aashish Rai, Srinath Sridhar

We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric v…