activity
20242026
most citedMoving Off-the-Grid: Scene-Grounded Video Representations

1 citations · 1 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CV2026

RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation

Peng Chen, Xiaobao Wei, Yi Yang +3

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limit…

cs.CV2025

2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model

Zilong Guo, Yi Luo, Long Sha +4

End-to-end autonomous driving has drawn tremendous attention recently. Many works focus on using modular deep neural networks to construct the end-to-end archi-tecture. However, wh…

cs.CV2025

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion

Yu Lu, Yi Yang

Recent advances in video generation models have enabled high-quality short video generation from text prompts. However, extending these models to longer videos remains a significan…

cs.CV2025

CHRIS: Clothed Human Reconstruction with Side View Consistency

Dong Liu, Yifan Yang, Zixiong Huang +2

Creating a realistic clothed human from a single-view RGB image is crucial for applications like mixed reality and filmmaking. Despite some progress in recent years, mainstream met…

cs.CV2025

TAPNext: Tracking Any Point (TAP) as Next Token Prediction

Artem Zholus, Carl Doersch, Yi Yang +7

Tracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction. Existing methods…

cs.CV2025

From Image to Video: An Empirical Study of Diffusion Representations

Pedro Vélez, Luisa F. Polanía, Yi Yang +4

Diffusion models have revolutionized generative modeling, enabling unprecedented realism in image and video synthesis. This success has sparked interest in leveraging their represe…