Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning
Andrew Li, Rahul Thapa, Rahul Chalamala +3
Vision-Language Models (VLMs) excel at understanding single images, aided by high-quality instruction datasets. However, multi-image reasoning remains underexplored in the open-sou…
cs.CV2024
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Rahul Thapa, Kezhen Chen, Ian Covert +4
Recent advances in vision-language models (VLMs) have demonstrated the advantages of processing images at higher resolutions and utilizing multi-crop features to preserve native re…