8 papers
Recurrent Video Masked Autoencoders
Daniel Zoran, Nikhil Parthasarathy, Yi Yang +3
We present Recurrent Video Masked-Autoencoders (RVM): a novel approach to video representation learning that leverages recurrent computation to model the temporal structure of vide…
RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation
Peng Chen, Xiaobao Wei, Yi Yang +3
Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limit…
2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model
Zilong Guo, Yi Luo, Long Sha +4
End-to-end autonomous driving has drawn tremendous attention recently. Many works focus on using modular deep neural networks to construct the end-to-end archi-tecture. However, wh…
Scaling 4D Representations
João Carreira, Dilara Gokay, Michael King +32
Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x20…
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
Yu Lu, Yi Yang
Recent advances in video generation models have enabled high-quality short video generation from text prompts. However, extending these models to longer videos remains a significan…
CHRIS: Clothed Human Reconstruction with Side View Consistency
Dong Liu, Yifan Yang, Zixiong Huang +2
Creating a realistic clothed human from a single-view RGB image is crucial for applications like mixed reality and filmmaking. Despite some progress in recent years, mainstream met…