3 papers
cs.CV2026
Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination
Ziqiang Shi, Rujie Liu, Shanshan Yu +2
Rapid progress in large vision-language models (LVLMs) has achieved unprecedented performance in vision-language tasks. However, due to the strong prior of large language models (L…
cs.CV2026
SchröMind: Mitigating Hallucinations in Multimodal Large Language Models via Solving the Schrödinger Bridge Problem
Ziqiang Shi, Rujie Liu, Shanshan Yu +2
Recent advancements in Multimodal Large Language Models (MLLMs) have achieved significant success across various domains. However, their use in high-stakes fields like healthcare r…
cs.CV2025
A Benchmark for Vision-Centric HD Mapping by V2I Systems
Miao Fan, Shanshan Yu, Shengtong Xu +3
Autonomous driving faces safety challenges due to a lack of global perspective and the semantic information of vectorized high-definition (HD) maps. Information from roadside camer…