Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
Rahul Thapa, Andrew Li, Qingyang Wu +8
Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these video…
cs.CV2024
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Rahul Thapa, Kezhen Chen, Ian Covert +4
Recent advances in vision-language models (VLMs) have demonstrated the advantages of processing images at higher resolutions and utilizing multi-crop features to preserve native re…