1 citations · 1 across the 2 of their papers we have counts for
9 papers
MR-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
Hong Jiang, Junnan Zhu, Jingwang Huang +9
Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information j…
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
Xiao Liu, Nayu Liu, Junnan Zhu +6
Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and complex question answering. Exis…
CoMMET: A Psychologically Grounded Benchmark for Evaluating Theory of Mind in Multimodal LLMs
Ruirui Chen, Weifeng Jiang, Chengwei Qin +3
Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Multimodal Large Language Models (MLLMs)…
Are Large Language Models Effective Knowledge Graph Constructors?
Ruirui Chen, Weifeng Jiang, Chengwei Qin +10
Knowledge graphs (KGs) are widely used in knowledge-intensive applications, yet it remains unclear how effectively current large language models (LLMs) can construct document-groun…
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
Yufei Gao, Jiaying Fei, Nuo Chen +4
Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings. We argue…
From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems
Xiuchao Sui, Daiying Tian, Qi Sun +4
Foundation models (FMs) are increasingly used to bridge language and action in embodied agents, yet the operational characteristics of different FM integration strategies remain un…