#chain-of-thought

topicchain-of-thought

8 papers · 1 filter

cs.CV2026

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

Shoutai Zhu, Tianyang Xu, Bin Sun +3

The paper introduces OPLD, an on‑policy latent distillation framework that transfers the reasoning ability of privileged multimodal chain‑of‑thought prompts into continuous latent…

cs.IR2026

WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models

Hao Jiang, Peiru Du, Pengfei Yao +10

The paper presents WhisperRec, a framework that compresses teacher-generated chain‑of‑thought explanations into learnable latent tokens, allowing recommendation models to reason in…

cs.CL2026

Scaling Evaluation-time Compute with Reasoning Models as Evaluators

Seungone Kim, Ian Wu, Jinu Lee +8

The paper studies how using larger, chain‑of‑thought reasoning language models as evaluators—by allocating more test‑time compute—can improve the accuracy of evaluating and reranki…

cs.LG2026

SCOPE-RL: Optimizing Reasoning Paths Before and After Success

Xiaojian Liu, Han Xu, Jianqiang Xia +6

The paper proposes SCOPE-RL, a two-stage reinforcement learning framework that adds dense, verifiable rewards to both pre‑success and post‑success reasoning steps of large language…

cs.CV2026

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection

Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2

The paper introduces EFlow, a framework that first retrieves visual evidence from long videos before reasoning, using separate chain‑of‑thought modules for temporal grounding and a…

cs.AI2026

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

Zeyu Chen, Huanjin Yao, Ziwang Zhao +1

The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…