18 citations · 20 across the 9 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
Rong Fan, Kaiyan Xiao, Minghao Zhu +3
Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However,…
cs.CV2024
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
Harshit, Tolga Tasdizen
The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Mode…
cs.CV2024
Towards Artwork Explanation in Large-scale Vision Language Models
Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2
Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…