2 papers
cs.CV2026
Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
Hongwei Fang, Jiahang Cai, Xun Wang +1
Vision Transformers (ViTs) have recently achieved state-of-the-art performance in 2D human pose estimation due to their strong global modeling capability. However, existing ViT-bas…
cs.CV2025
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
Yonghui Yu, Jiahang Cai, Xun Wang +1
Existing multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single person pose…