Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
Yuan Zang, Hao Tan, Seunghyun Yoon +5
We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames.…
cs.CV2024
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
Yuan Zang, Tian Yun, Hao Tan +2
Do vision-language models (VLMs) pre-trained to caption an image of a "durian" learn visual concepts such as "brown" (color) and "spiky" (texture) at the same time? We aim to answe…