2 papers
cs.CV2025
Locality Alignment Improves Vision-Language Models
Ian Covert, Tony Sun, James Zou +1
Vision language models (VLMs) have seen growing adoption in recent years, but many still struggle with basic spatial reasoning errors. We hypothesize that this is due to VLMs adopt…
cs.CV2024
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Rahul Thapa, Kezhen Chen, Ian Covert +4
Recent advances in vision-language models (VLMs) have demonstrated the advantages of processing images at higher resolutions and utilizing multi-crop features to preserve native re…