2 papers
cs.CV2025
DropMAE: Learning Representations via Masked Autoencoders with Spatial-Attention Dropout for Temporal Matching Tasks
Qiangqiang Wu, Tianyu Yang, Ziquan Liu +3
This paper studies masked autoencoder (MAE) video pre-training for various temporal matching-based downstream tasks, i.e., object-level tracking tasks including video object tracki…
cs.CV2024
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
Qi Zhang, Antoni B. Chan
Recently multi-view crowd counting using deep neural networks has been proposed to enable counting in large and wide scenes using multiple cameras. The current methods project the…