3 papers
cs.CV2026
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
Yuxiang Ji, Yong Wang, Ziyu Ma +6
The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches le…
cs.AI2025
Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
Ziyu Ma, Chenhui Gou, Yiming Hu +4
Large Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to many-shot settings remains difficult due to limited context length and hi…
cs.CV2025
An Empirical Study on How Video-LLMs Answer Video Questions
Chenhui Gou, Ziyu Ma, Zicheng Duan +6
Taking advantage of large-scale data and pretrained language models, Video Large Language Models (Video-LLMs) have shown strong capabilities in answering video questions. However,…