Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
Tengda Guo, Jie Leng, Hanlei Li +6
Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual information to support semantic…
cs.CV2024
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
Jinyuan Li, Han Li, Di Sun +4
Grounded Multimodal Named Entity Recognition (GMNER) is a nascent multimodal task that aims to identify named entities, entity types and their corresponding visual regions. GMNER t…