2 papers
cs.CL2026
Guided Verifier: Collaborative Multimodal Reasoning via Dynamic Process Supervision
Lingzhuang Sun, Ruitong Liu, Yuxia Zhu +5
Reinforcement Learning (RL) has emerged as a pivotal mechanism for enhancing the complex reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevailing par…
cs.LG2025
Pinpointing crucial steps: Attribution-based Credit Assignment for Verifiable Reinforcement Learning
Junxi Yin, Haisen Luo, Zhenyu Li +4
While Reinforcement Learning with Verifiable Rewards (RLVR) enhances complex reasoning in LLMs, current methods struggle to balance exploration and exploitation. This leads to crit…