2 papers
cs.CV2026
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
Tengda Guo, Jie Leng, Hanlei Li +6
Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual information to support semantic…
cs.MM2025
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
Jinyuan Li, Ziyan Li, Han Li +4
Grounded Multimodal Named Entity Recognition (GMNER) task aims to identify named entities, entity types and their corresponding visual regions. GMNER task exhibits two challenging…