3 papers
cs.CV2025
Distributed Zero-Shot Learning for Visual Recognition
Zhi Chen, Yadan Luo, Zi Huang +3
In this paper, we propose a Distributed Zero-Shot Learning (DistZSL) framework that can fully exploit decentralized data to learn an effective model for unseen classes. Considering…
cs.CL2025
GeoExplain: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
Fenghua Cheng, Jinxiang Wang, Sen Wang +2
Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention. Al…
cs.CV2025
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
Zixin Wang, Dong Gong, Sen Wang +2
Contrastive Language-Image Pretraining (CLIP) excels at learning generalizable image representations but often falls short in zero-shot inference on certain downstream datasets. Te…