#multimodal reasoning
20 papers · 1 filter
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
Enjun Du, Hange Zhou, Chenxu Du +4
The paper introduces LedgerMind, a framework that records and constrains the evidence used by multimodal agents during visual question answering, ensuring that each reasoning step…
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
Feng Xiong, Leyan Xue, Hongyu Lin
The paper proposes Perception-Correction Distillation (PCD), a label‑free method that uses downstream failures and teacher‑student disagreement to pinpoint and correct perception e…
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
Haoqing Wang, Xingrun Xing, Wei Xia +2
FaithEyes proposes a multi‑agent framework where a vision‑language model judges its own tool calls to ensure they are useful, improving both accuracy and tool faithfulness on visua…
RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
Jingxiang Fan, Junbao Zhuo, Bochao Zou
The paper proposes Reflective Retrieval Memory (RRM), a framework that adds a reflective experience memory to an entity‑centric multimodal memory graph, enabling agents to learn an…
OPLD: On-Policy Latent Distillation for Multimodal Reasoning
Shoutai Zhu, Tianyang Xu, Bin Sun +3
The paper introduces OPLD, an on‑policy latent distillation framework that transfers the reasoning ability of privileged multimodal chain‑of‑thought prompts into continuous latent…
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
Shiwei Gan, Lichen Wang, Xiao Liu +4
The paper introduces Sign Language Question Answering (SLQA), a task where models answer natural language questions about sign language videos, and provides two benchmark datasets…