Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
Anas Anwarul Haq Khan, Mariam Husain, Pratik Jalan +1
Vision-language pretraining has driven much of the recent progress in medical image representation learning, but this paradigm is constrained by the availability of paired image-te…
cs.CV2025
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
Anas Anwarul Haq Khan, Utkarsh Verma, Ganesh Ramakrishnan
We introduce DEEVISum (Distilled Early Exit Vision language model for Summarization), a lightweight, efficient, and scalable vision language model designed for segment wise video s…
cs.CV2024
Sumotosima: A Framework and Dataset for Classifying and Summarizing Otoscopic Images
Eram Anwarul Khan, Anas Anwarul Haq Khan
Otoscopy is a diagnostic procedure to examine the ear canal and eardrum using an otoscope. It identifies conditions like infections, foreign bodies, ear drum perforations and ear a…