7 papers
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
Yongxin Wang, Ruizhe Zhou, Yueling Tang +4
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text,…
MOGeo: Beyond One-to-One Cross-View Object Geo-localization
Bo Lv, Qingwang Zhang, Le Wu +2
Cross-View Object Geo-Localization (CVOGL) aims to locate an object of interest in a query image within a corresponding satellite image. Existing methods typically assume that the…
From Horizontal to Rotated: Cross-View Object Geo-Localization with Orientation Awareness
Chenlin Fu, Ao Gong, Yingying Zhu
Cross-View object geo-localization (CVOGL) aims to precisely determine the geographic coordinates of a query object from a ground or drone perspective by referencing a satellite ma…
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
Yongxin Wang, Zhicheng Yang, Meng Cao +5
Group-relative reinforcement learning with verifiable rewards (RLVR) often wastes the most informative data it already has the failures. When all rollouts are wrong, gradients stal…
Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization
Shuhan Hu, Yiru Li, Yuanyuan Li +1
Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and d…
SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval
Yuqi Xiao, Yingying Zhu
Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve a target image given a reference image and a relative text, without relying on costly triplet annotations. Existing CLI…