3 papers
cs.CV2026
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
Brandon Collins, Logan Bolton, Hung Huy Nguyen +3
When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro an…
cs.CV2025
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
Yuan Zang, Hao Tan, Seunghyun Yoon +5
We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames.…
cs.CV2025
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
Yuan Zang, Tian Yun, Hao Tan +2
Do vision-language models (VLMs) pre-trained to caption an image of a "durian" learn visual concepts such as "brown" (color) and "spiky" (texture) at the same time? We aim to answe…