5 citations · 8 across the 16 of their papers we have counts for
Showing 2026 · cs.CVShow all
2 papers · 2 filters
cs.CV2026
Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models
Zhiwei Yang, Yuanchen Wu, Nan Zhang +3
Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated objects and neglect structu…
cs.CV2026
Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
Enguang Wang, Qiang Wang, Yuanchen Wu +5
While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear…