3 papers
cs.CV2026
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
Brandon Collins, Logan Bolton, Hung Huy Nguyen +3
When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro an…
cs.CV2025
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
Hung Huy Nguyen, Pooyan Rahmanzadehgervi, Long Mai +1
Detecting object-level changes between two images across possibly different views is a core task in many applications that involve visual inspection or camera surveillance. Existin…
cs.CV2024
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu +2
Multi-head self-attention (MHSA) is a key component of Transformers, a widely popular architecture in both language and vision. Multiple heads intuitively enable different parallel…