9 papers
One-Shot Crowd Counting With Density Guidance For Scene Adaptation
Jiwei Chen, Qi Wang, Junyu Gao +3
Crowd scenes captured by cameras at different locations vary greatly, and existing crowd models have limited generalization for unseen surveillance scenes. To improve the generaliz…
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
Zunhai Su, Weihao Ye, Hansen Feng +5
Learning-based 3D visual geometry models have significantly advanced with the advent of large-scale transformers. Among these, StreamVGGT leverages frame-wise causal attention to d…
CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment
Yunzuo Hu, Wen Li, Jing Zhang
Audio-visual learning suffers from modality misalignment caused by off-screen sources and background clutter, and current methods usually amplify irrelevant regions or moments, lea…
Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
Zhong Peng, Yishi Xu, Gerong Wang +4
Compositional Zero-Shot Learning (CZSL) seeks to recognize unseen state-object pairs by recombining primitives learned from seen compositions. Despite recent progress with vision-l…
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
Yuan Zhou, Qingshan Xu, Jiequan Cui +4
Recently, large efforts have been made to design efficient linear-complexity visual Transformers. However, current linear attention models are generally unsuitable to be deployed i…
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
Tianran Ouyang, Xingping Dong, Jing Zhang +3
Slot Attention, an approach that binds different objects in a scene to a set of "slots", has become a leading method in unsupervised object-centric learning. Most methods assume a…