1 citations · 1 across the 1 of their papers we have counts for
3 papers · 1 filter
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
Anas Anwarul Haq Khan, Mariam Husain, Pratik Jalan +1
Vision-language pretraining has driven progress in medical image representation learning, but it depends on paired image-text data and can inherit reporting bias from clinical narr…
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
Anas Anwarul Haq Khan, Utkarsh Verma, Ganesh Ramakrishnan
We introduce DEEVISum (Distilled Early Exit Vision language model for Summarization), a lightweight, efficient, and scalable vision language model designed for segment wise video s…
Sumotosima: A Framework and Dataset for Classifying and Summarizing Otoscopic Images
Eram Anwarul Khan, Anas Anwarul Haq Khan
Otoscopy is a diagnostic procedure to examine the ear canal and eardrum using an otoscope. It identifies conditions like infections, foreign bodies, ear drum perforations and ear a…