2 papers
cs.CV2026
Modular Diffusion Models for Structured Visual Recognition
Siddhesh Khandelwal, Björn Ommer, Leonid Sigal
Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often produce deterministic, fixed o…
cs.CV2024
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
Jiayun Luo, Siddhesh Khandelwal, Leonid Sigal +1
From image-text pairs, large-scale vision-language models (VLMs) learn to implicitly associate image regions with words, which prove effective for tasks like visual question answer…