Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
Rahul Thapa, Andrew Li, Qingyang Wu +8
Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these video…
cs.CV2025
SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning
Andrew Li, Rahul Thapa, Rahul Chalamala +3
Vision-Language Models (VLMs) excel at understanding single images, aided by high-quality instruction datasets. However, multi-image reasoning remains underexplored in the open-sou…