68 citations · 98 across the 18 of their papers we have counts for
27 papers
DB-SAM: Delving into High Quality Universal Medical Image Segmentation
Chao Qin, Jiale Cao, Huazhu Fu +2
Recently, the Segment Anything Model (SAM) has demonstrated promising segmentation capabilities in a variety of downstream segmentation tasks. However in the context of universal m…
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
Umair Nawaz, Muhammad Awais, Hanan Gani +4
Capitalizing on vast amount of image-text data, large-scale vision-language pre-training has demonstrated remarkable zero-shot capabilities and has been utilized in several applica…
CDChat: A Large Multimodal Model for Remote Sensing Change Description
Mubashir Noman, Noor Ahsan, Muzammal Naseer +4
Large multimodal models (LMMs) have shown encouraging performance in the natural image domain using visual instruction tuning. However, these LMMs struggle to describe the content…
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
Asif Hanif, Fahad Shamshad, Muhammad Awais +5
Medical foundation models are gaining prominence in the medical community for their ability to derive general representations from extensive collections of medical image-text pairs…
Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning
Amandeep Kumar, Muhammad Awais, Sanath Narayan +3
Drawing upon StyleGAN's expressivity and disentangled latent space, existing 2D approaches employ textual prompting to edit facial images with different attributes. In contrast, 3D…
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
Omkar Thawakar, Muzammal Naseer, Rao Muhammad Anwer +4
Composed video retrieval (CoVR) is a challenging problem in computer vision which has recently highlighted the integration of modification text with visual queries for more sophist…