2 citations · 2 across the 10 of their papers we have counts for
13 papers · 1 filter
HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control
Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu +3
We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand ann…
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv +5
Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horiz…
SemanticGen: Video Generation in Semantic Space
Jianhong Bai, Xiaoshi Wu, Xintao Wang +9
State-of-the-art video generative models typically learn the distribution of video latents in the VAE space and map them to pixels using a VAE decoder. While this approach can gene…
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
Minglei Shi, Haolin Wang, Borui Zhang +11
Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generati…
Latent Diffusion Model without Variational Autoencoder
Minglei Shi, Haolin Wang, Wenzhao Zheng +6
Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with variational autoencoders (VAEs). While effective for high-fidelity synthesis…
HPSv3: Towards Wide-Spectrum Human Preference Score
Yuhang Ma, Yunhao Shui, Xiaoshi Wu +2
Evaluating text-to-image generation models requires alignment with human perception, yet existing human-centric metrics are constrained by limited data coverage, suboptimal feature…