3 papers
cs.CV2025
Distribution Matching Variational AutoEncoder
Sen Ye, Jianning Pei, Mengde Xu +4
Most visual generative models compress images into a latent space before applying diffusion or autoregressive modelling. Yet, existing approaches such as VAEs and foundation model…
cs.CV2023
Multiple View Geometry Transformers for 3D Human Pose Estimation
Ziwei Liao, Jialiang Zhu, Chunyu Wang +2
In this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer…
cs.CV2023
GAIA: Zero-shot Talking Avatar Generation
Tianyu He, Junliang Guo, Runyi Yu +10
Zero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics…