Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation
Xinyan Ye, Jiankang Deng, Abbas Edalat
Talking-head generation requires joint modeling of identity, head pose, facial expression, and mouth dynamics. Existing methods typically address only a subset of these factors, an…
cs.CV2023
STDiff: Spatio-temporal Diffusion for Continuous Stochastic Video Prediction
Xi Ye, Guillaume-Alexandre Bilodeau
Predicting future frames of a video is challenging because it is difficult to learn the uncertainty of the underlying factors influencing their contents. In this paper, we propose…
cs.CV2022
VPTR: Efficient Transformers for Video Prediction
Xi Ye, Guillaume-Alexandre Bilodeau
In this paper, we propose a new Transformer block for video future frames prediction based on an efficient local spatial-temporal separation attention mechanism. Based on this new…