241 citations · 330 across the 13 of their papers we have counts for
15 papers · 1 filter
Mind the Discriminability Trap in Source-Free Cross-domain Few-shot Learning
Zhenyu Zhang, Yixiong Zou, Yuhua Li +2
Source-Free Cross-Domain Few-Shot Learning (SF-CDFSL) focuses on fine-tuning with limited training data from target domains (e.g., medical or satellite images), where Vision-Langua…
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
Baiyang Song, Jun Peng, Yuxin Zhang +3
Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static fra…
Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot Learning
Zhenyu Zhang, Guangyao Chen, Yixiong Zou +2
The Contrastive Language-Image Pre-Training (CLIP) model excels in few-shot learning by aligning visual and textual representations. Our study shows that template-sample similarity…
Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding
Qingyang Yan, Guangyao Chen, Yixiong Zou
Chain-of-Thought (CoT) prompting has recently shown significant promise across various NLP and computer vision tasks by explicitly generating intermediate reasoning steps. However,…
When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid Network
Dong Xiao, Guangyao Chen, Peixi Peng +4
Anomaly detection is essential for the safety and reliability of autonomous driving systems. Current methods often focus on detection accuracy but neglect response time, which is c…
Adapter Naturally Serves as Decoupler for Cross-Domain Few-Shot Semantic Segmentation
Jintao Tong, Ran Ma, Yixiong Zou +3
Cross-domain few-shot segmentation (CD-FSS) is proposed to pre-train the model on a source-domain dataset with sufficient samples, and then transfer the model to target-domain data…