#multimodal reasoning

topicmultimodal reasoning

20 papers · 1 filter

cs.LG2026

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

Enjun Du, Hange Zhou, Chenxu Du +4

The paper introduces LedgerMind, a framework that records and constrains the evidence used by multimodal agents during visual question answering, ensuring that each reasoning step…

cs.AI2026

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners

Feng Xiong, Leyan Xue, Hongyu Lin

The paper proposes Perception-Correction Distillation (PCD), a label‑free method that uses downstream failures and teacher‑student disagreement to pinpoint and correct perception e…

cs.CV2026

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

Haoqing Wang, Xingrun Xing, Wei Xia +2

FaithEyes proposes a multi‑agent framework where a vision‑language model judges its own tool calls to ensure they are useful, improving both accuracy and tool faithfulness on visua…

cs.CL2026

RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning

Jingxiang Fan, Junbao Zhuo, Bochao Zou

The paper proposes Reflective Retrieval Memory (RRM), a framework that adds a reflective experience memory to an entity‑centric multimodal memory graph, enabling agents to learn an…

cs.CV2026

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

Shoutai Zhu, Tianyang Xu, Bin Sun +3

The paper introduces OPLD, an on‑policy latent distillation framework that transfers the reasoning ability of privileged multimodal chain‑of‑thought prompts into continuous latent…

cs.AI2026

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

Shiwei Gan, Lichen Wang, Xiao Liu +4

The paper introduces Sign Language Question Answering (SLQA), a task where models answer natural language questions about sign language videos, and provides two benchmark datasets…