11 papers · 1 filter
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
Jing Zhang, Wentao Jiang, Tao Huang +8
Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: special…
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
Yi Liu, Jing Zhang, Di Wang +3
Multimodal large language models (MLLMs) suffer from pronounced hallucinations in remote sensing visual question-answering (RS-VQA), primarily caused by visual grounding failures i…
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
Fengxiang Wang, Mingshuo Chen, Yueying Li +10
The "thinking-with-images" paradigm enables multimodal large language models (MLLMs) to actively explore visual scenes via zoom-in tools. This is essential for ultra-high-resolutio…
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
Tianran Ouyang, Xingping Dong, Jing Zhang +3
Slot Attention, an approach that binds different objects in a scene to a set of "slots", has become a leading method in unsupervised object-centric learning. Most methods assume a…
GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
Zixuan Song, Jing Zhang, Di Wang +5
Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradi…
GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
Di Wang, Shunyu Liu, Wentao Jiang +10
Multimodal large language models (MLLMs) have undergone rapid development in advancing geospatial scene understanding. Recent studies have sought to enhance the reasoning capabilit…