14 papers · 1 filter
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
Xike Zhang, Maoyuan Ye, Juhua Liu +1
Previous works based on Segment Anything Model (SAM) have achieved promising performance in unified scene text detection and layout analysis. However, the typical reliance on pixel…
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
Shuo Ni, Tong Wang, Jing Zhang +4
Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scale mismatch between large-sca…
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
Qihuang Zhong, Liang Ding, Wenjie Xuan +3
Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acquiring high-quality reasoning…
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
Haibin He, Maoyuan Ye, Jing Zhang +2
Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despite the strong multimodal vide…
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
Xiao He, Huangxuan Zhao, Guojia Wan +8
Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and u…
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
Haibin He, Jing Zhang, Maoyuan Ye +3
Video text spotting (VTS) extends image text spotting (ITS) by adding text tracking, significantly increasing task complexity. Despite progress in VTS, existing methods still fall…