Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards
Yupeng Chang, Yuan Wu, Yi Chang
Critic-free reinforcement learning with verifiable rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), avoids training a value function (critic) and reduces m…
cs.AI2025
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
Yue Zhou, Yi Chang, Yuan Wu
Reasoning is a critical capability of multimodal large language models (MLLMs) for solving complex multimodal tasks, and judging the correctness of reasoning steps is crucial for i…
cs.AI2025
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework
Jialin Li, Jinzhe Li, Gengxu Li +2
With the advancement of code generation capabilities in large language models (LLMs), their reliance on input premises has intensified. When users provide inputs containing faulty…