Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping
Rishubh Parihar, Ayush Raina, R. Venkatesh Babu +1
Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis. However, these models are co…
cs.CV2026
The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
Kelly Cui, Nikhil Prakash, Shoval Messica +4
Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Ye…