3 papers
cs.CV2025
Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
Peter Robicheaux, Matvei Popov, Anish Madan +4
Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of…
cs.CV2025
SMORE: Simultaneous Map and Object REconstruction
Nathaniel Chodosh, Anish Madan, Simon Lucey +1
We present a method for dynamic surface reconstruction of large-scale urban scenes from LiDAR. Depth-based reconstructions tend to focus on small-scale objects or large-scale SLAM…
cs.CV2024
Revisiting Few-Shot Object Detection with Vision-Language Models
Anish Madan, Neehar Peri, Shu Kong +1
The era of vision-language models (VLMs) trained on web-scale datasets challenges conventional formulations of "open-world" perception. In this work, we revisit the task of few-sho…