2 papers
cs.CV2024
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
Wangyu Xue, Chen Qian, Jiayi Wu +5
Existing works on human-centric video understanding typically focus on analyzing specific moment or entire videos. However, many applications require higher precision at the frame…
cs.CV2024
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
Tianfang Zhang, Lei Li, Yang Zhou +4
Vision Transformers (ViTs) mark a revolutionary advance in neural networks with their token mixer's powerful global context capability. However, the pairwise token affinity and com…