3 papers
cs.LG2025
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
Sunil Kumar, Bowen Zhao, Leo Dirac +1
Despite tremendous recent advances in large model reasoning ability, vision-language models (VLMs) still struggle with detailed visual reasoning, especially when compute resources…
cs.LG2024
Fine-tuning Vision Classifiers On A Budget
Sunil Kumar, Ted Sandler, Paulina Varshavskaya
Fine-tuning modern computer vision models requires accurately labeled data for which the ground truth may not exist, but a set of multiple labels can be obtained from labelers of v…
cs.CV2024
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
Bowen Zhao, Leo Parker Dirac, Paulina Varshavskaya
Large vision-language models (VLMs) have become state-of-the-art for many computer vision tasks, with in-context learning (ICL) as a popular adaptation strategy for new ones. But c…