5 papers
HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control
Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu +3
We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand ann…
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv +5
Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horiz…
Latent Diffusion Model without Variational Autoencoder
Minglei Shi, Haolin Wang, Wenzhao Zheng +6
Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with variational autoencoders (VAEs). While effective for high-fidelity synthesis…
SemanticGen: Video Generation in Semantic Space
Jianhong Bai, Xiaoshi Wu, Xintao Wang +9
State-of-the-art video generative models typically learn the distribution of video latents in the VAE space and map them to pixels using a VAE decoder. While this approach can gene…
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
Minglei Shi, Haolin Wang, Borui Zhang +11
Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generati…