activity
20242026
most citedCausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes

2 citations · 2 across the 2 of their papers we have counts for

collaborators

9 papers

cs.CL2026

MR-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

Hong Jiang, Junnan Zhu, Jingwang Huang +9

Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information j…

cs.CV20262 cited

CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes

Paritosh Parmar, Eric Peh, Ruirui Chen +4

Causal video question answering (QA) has garnered increasing interest, yet existing datasets often lack depth in causal reasoning. To address this gap, we capitalize on the unique…

cs.CV2026

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding

Xiao Liu, Nayu Liu, Junnan Zhu +6

Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and complex question answering. Exis…

cs.CL2026

CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?

Ruirui Chen, Weifeng Jiang, Chengwei Qin +1

Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Large Language Models (LLMs) become ubiqu…

cs.RO2025

From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems

Xiuchao Sui, Daiying Tian, Qi Sun +4

Foundation models (FMs) are increasingly used to bridge language and action in embodied agents, yet the operational characteristics of different FM integration strategies remain un…

cs.CL2025

Are Large Language Models Effective Knowledge Graph Constructors?

Ruirui Chen, Weifeng Jiang, Chengwei Qin +4

Knowledge graphs (KGs) are vital for knowledge-intensive tasks and have shown promise in reducing hallucinations in large language models (LLMs). However, constructing high-quality…