3 papers
cs.CL2026
MR-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
Hong Jiang, Junnan Zhu, Jingwang Huang +9
Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information j…
cs.CV2026
LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents
Zijian Wang, Junnan Zhu, Rongzhen Li +7
Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this ch…
cs.CV2026
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
Xiao Liu, Nayu Liu, Junnan Zhu +6
Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and complex question answering. Exis…