Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
Logan Lawrence, Oindrila Saha, Megan Wei +3
Despite the renewed interest in zero-shot visual classification due to the rise of Multimodal Large Language Models (MLLMs), the problem of evaluating free-form responses of auto-r…
cs.CV2025
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
Yuan Zang, Hao Tan, Seunghyun Yoon +5
We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames.…