4 papers · 1 filter
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
Sathira Silva, Eman Ali, Chetan Arora +1
Unsupervised adaptation of CLIP-based vision-language models (VLMs) for fine-grained image classification requires sensitivity to microscopic local cues. While CLIP exhibits strong…
Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
Eman Ali, Sathira Silva, Chetan Arora +1
Vision-language models (VLMs) like CLIP excel in zero-shot learning by aligning image and text representations through contrastive pretraining. Existing approaches to unsupervised…
DPA: Dual Prototypes Alignment for Unsupervised Adaptation of Vision-Language Models
Eman Ali, Sathira Silva, Muhammad Haris Khan
Vision-language models (VLMs), e.g., CLIP, have shown remarkable potential in zero-shot image classification. However, adapting these models to new domains remains challenging, esp…
Noise-Tolerant Few-Shot Unsupervised Adapter for Vision-Language Models
Eman Ali, Muhammad Haris Khan
Recent advances in large-scale vision-language models have achieved impressive performance in various zero-shot image classification tasks. While prior studies have demonstrated si…