3 papers
cs.CV2026
Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case
Emilie Durrieu, Christophe Hurter, Philippe Muller +1
When humans play geolocation games such as GeoGuessr, they rely on concrete visual cues, such as road markings, vegetation, or architectural details, to infer where an image was ca…
cs.CL2026
Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
Galann Pennec, Zhengyuan Liu, Nicholas Asher +2
Vision-Language Models (VLMs) are able to process increasingly longer videos. Yet, important visual information is easily lost throughout the entire context and missed by VLMs. Als…
cs.CL2025
Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
Galann Pennec, Zhengyuan Liu, Nicholas Asher +2
Vision-Language Models (VLMs) often struggle to balance visual and textual information when summarizing complex multimodal inputs, such as entire TV show episodes. In this paper, w…