1 citations · 1 across the 13 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
Shezheng Song, Shasha Li, Jie Yu
Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a variety of vision-language tasks. However, their internal reasoning often exhibits a critica…
cs.CV2026
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
Shezheng Song, Shasha Li, Shan Zhao +6
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet how they internally integrate visual and textual information remain…
cs.CV2024★ 1 cited
MOSABench: Multi-Object Sentiment Analysis Benchmark for Evaluating Multimodal Large Language Models Understanding of Complex Image
Shezheng Song, Chengxiang He, Shan Zhao +4
Multimodal large language models (MLLMs) have shown remarkable progress in high-level semantic tasks such as visual question answering, image captioning, and emotion recognition. H…