Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
Yoonjeon Kim, Doohyuk Jang, Eunho Yang
Recent research on reasoning models explores the meta-awareness of language models, including their ability to determine optimal thinking duration, recognize knowledge boundaries,…
cs.LG2026
Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards
Haechan Kim, Soohyun Ryu, Gyouk Chu +2
Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective post-training paradigm for improving the reasoning capabilities of large language models. However,…