collaborators

9 papers

cs.CV2026

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

Haoyu Yang, Meixing Shi, Zengjie Chen +5

Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack pre…

cs.CV2026

EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation

Chengyi Peng, Haoyu Yang, Meixing Shi +2

Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, prior,uncertain, or irrelevant findings, whi…

cs.CV2026

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

Bin Chen, Yuxiang Cai, Yadan Luo +3

Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performance. Existing token pruning m…

cs.LG2026

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding

Yingxuan Zhuang, Jingxiao Yang, Miao Pan +7

MLLMs frequently hallucinate objects inconsistent with visual inputs. This issue is typically attributed to the over-reliance on language priors, which can override the visual cont…

cs.CV2026

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

Qiaoru Li, Shaotian Liang, Jintao Chen +4

Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead of chain-of-thought for medica…

cs.CV2026

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

Zilun Zhang, Zian Guan, Tiancheng Zhao +7

Referring expression understanding in remote sensing poses unique challenges, as it requires reasoning over complex object-context relationships. While supervised fine-tuning (SFT)…