#chain-of-thought
8 papers · 1 filter
OPLD: On-Policy Latent Distillation for Multimodal Reasoning
Shoutai Zhu, Tianyang Xu, Bin Sun +3
The paper introduces OPLD, an on‑policy latent distillation framework that transfers the reasoning ability of privileged multimodal chain‑of‑thought prompts into continuous latent…
WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models
Hao Jiang, Peiru Du, Pengfei Yao +10
The paper presents WhisperRec, a framework that compresses teacher-generated chain‑of‑thought explanations into learnable latent tokens, allowing recommendation models to reason in…
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
Seungone Kim, Ian Wu, Jinu Lee +8
The paper studies how using larger, chain‑of‑thought reasoning language models as evaluators—by allocating more test‑time compute—can improve the accuracy of evaluating and reranki…
SCOPE-RL: Optimizing Reasoning Paths Before and After Success
Xiaojian Liu, Han Xu, Jianqiang Xia +6
The paper proposes SCOPE-RL, a two-stage reinforcement learning framework that adds dense, verifiable rewards to both pre‑success and post‑success reasoning steps of large language…
EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection
Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2
The paper introduces EFlow, a framework that first retrieves visual evidence from long videos before reasoning, using separate chain‑of‑thought modules for temporal grounding and a…
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Zeyu Chen, Huanjin Yao, Ziwang Zhao +1
The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…