8 papers · 1 filter
Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models
Yunkai Dang, Yifan Jiang, Yizhu Jiang +3
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ensuring their reliability in p…
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing
Yunkai Dang, Minxin Dai, Yuekun Yang +4
Ultra-high-resolution (UHR) remote sensing imagery couples kilometer-scale context with query-critical evidence that may occupy only a few pixels. Such vast spatial scale leads to…
CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models
Yunkai Dang, Yizhu Jiang, Yifan Jiang +4
Multimodal Large Language Models (MLLMs) suffer from substantial computational overhead due to the high redundancy in visual token sequences. Existing approaches typically address…
Prompt-Free Universal Region Proposal Network
Qihong Tang, Changhan Liu, Shaofeng Zhang +3
Identifying potential objects is critical for object recognition and analysis across various computer vision applications. Existing methods typically localize potential objects by…
FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing
Yunkai Dang, Donghao Wang, Jiacheng Yang +7
Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain…
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
Yunkai Dang, Meiyi Zhu, Donghao Wang +7
Multimodal large language models (MLLMs) demonstrate strong perception and reasoning performance on existing remote sensing (RS) benchmarks. However, most prior benchmarks rely on…