2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 2 cited
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
Umair Nawaz, Muhammad Awais, Hanan Gani +4
Capitalizing on vast amount of image-text data, large-scale vision-language pre-training has demonstrated remarkable zero-shot capabilities and has been utilized in several applica…
cs.CV2024★ 1 cited
CDChat: A Large Multimodal Model for Remote Sensing Change Description
Mubashir Noman, Noor Ahsan, Muzammal Naseer +4
Large multimodal models (LMMs) have shown encouraging performance in the natural image domain using visual instruction tuning. However, these LMMs struggle to describe the content…