4 papers
A Proposal-Free Query-Guided Network for Grounded Multimodal Named Entity Recognition
Hongbing Li, Jiamin Liu, Shuo Zhang +1
Grounded Multimodal Named Entity Recognition (GMNER) identifies named entities, including their spans and types, in natural language text and grounds them to the corresponding regi…
BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding
Hongbing Li, Linhui Xiao, Zihan Zhao +4
Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent…
Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models
Jingming Liu, Yumeng Li, Boyuan Xiao +5
Under pure textual modality, Large Language Models (LLMs) have demonstrated remarkable success in complex reasoning tasks by decomposing them into simpler sub-problems. However, Mu…
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
Yi Li, Haonan Wang, Qixiang Zhang +4
The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while…