3 papers
cs.CV2026
Video Generation with Predictive Latents
Yian Zhao, Feng Wang, Qiushan Guo +4
Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, improving training efficiency an…
cs.CV2024
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
Peng Jin, Hao Li, Zesen Cheng +6
Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global mot…
cs.CV2024
Technique Report of CVPR 2024 PBDL Challenges
Ying Fu, Yu Li, Shaodi You +96
The intersection of physics-based vision and deep learning presents an exciting frontier for advancing computer vision technologies. By leveraging the principles of physics to info…