Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
Qiaoru Li, Shaotian Liang, Jintao Chen +4
Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead of chain-of-thought for medica…
cs.CV2026
Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization
Miao Pan, Wangjie Gan, Jintao Chen +4
While Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse tasks, their practical deployment is severely hindered by hallucination issues, which…
cs.CV2025
TK-Mamba: Marrying KAN With Mamba for Text-Driven 3D Medical Image Segmentation
Haoyu Yang, Yutong Guan, Meixing Shi +8
3D medical image segmentation is important for clinical diagnosis and treatment but faces challenges from high-dimensional data and complex spatial dependencies. Traditional single…