2 citations · 2 across the 2 of their papers we have counts for
9 papers
MR-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
Hong Jiang, Junnan Zhu, Jingwang Huang +9
Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information j…
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
Paritosh Parmar, Eric Peh, Ruirui Chen +4
Causal video question answering (QA) has garnered increasing interest, yet existing datasets often lack depth in causal reasoning. To address this gap, we capitalize on the unique…
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
Xiao Liu, Nayu Liu, Junnan Zhu +6
Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and complex question answering. Exis…
CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?
Ruirui Chen, Weifeng Jiang, Chengwei Qin +1
Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Large Language Models (LLMs) become ubiqu…
From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems
Xiuchao Sui, Daiying Tian, Qi Sun +4
Foundation models (FMs) are increasingly used to bridge language and action in embodied agents, yet the operational characteristics of different FM integration strategies remain un…
Are Large Language Models Effective Knowledge Graph Constructors?
Ruirui Chen, Weifeng Jiang, Chengwei Qin +4
Knowledge graphs (KGs) are vital for knowledge-intensive tasks and have shown promise in reducing hallucinations in large language models (LLMs). However, constructing high-quality…