9 papers
Mind the Discriminability Trap in Source-Free Cross-domain Few-shot Learning
Zhenyu Zhang, Yixiong Zou, Yuhua Li +2
Source-Free Cross-Domain Few-Shot Learning (SF-CDFSL) focuses on fine-tuning with limited training data from target domains (e.g., medical or satellite images), where Vision-Langua…
Reclaiming Lost Text Layers for Source-Free Cross-Domain Few-Shot Learning
Zhenyu Zhang, Guangyao Chen, Yixiong Zou +2
Source-Free Cross-Domain Few-Shot Learning (SF-CDFSL) focuses on fine-tuning with limited training data from target domains (e.g., medical or satellite images), where CLIP has rece…
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
Baiyang Song, Jun Peng, Yuxin Zhang +3
Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static fra…
Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot Learning
Zhenyu Zhang, Guangyao Chen, Yixiong Zou +2
The Contrastive Language-Image Pre-Training (CLIP) model excels in few-shot learning by aligning visual and textual representations. Our study shows that template-sample similarity…
Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding
Qingyang Yan, Guangyao Chen, Yixiong Zou
Chain-of-Thought (CoT) prompting has recently shown significant promise across various NLP and computer vision tasks by explicitly generating intermediate reasoning steps. However,…
From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors via LLM-guided Symbolic Reasoning
Yuhui Zeng, Haoxiang Wu, Wenjie Nie +6
Current object detectors excel at entity localization and classification, yet exhibit inherent limitations in event recognition capabilities. This deficiency arises from their arch…