Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
Bo Yu, Fengze Yang, Yiming Liu +6
The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven infe…
cs.CV2026
FreeFly-Thinking : Aligning Chain-of-Thought Reasoning with Continuous UAV Navigation
Jiaxu Zhou, Shaobo Wang, Zhiyuan Yang +2
Vision-Language Navigation aims to enable agents to understand natural language instructions and carry out appropriate navigation actions in real-world environments. Most work focu…
cs.CV2024
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
Tao Li, Linjun Shou, Xuejun Liu
Zero-shot visual question answering (VQA) is a challenging task that requires reasoning across modalities. While some existing methods rely on a single rationale within the Chain o…