3 papers
cs.LG2025
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
Satwik Bhattamishra, Phil Blunsom, Varun Kanade
We study the learnability of languages in the Next Symbol Prediction (NSP) setting, where a learner receives only positive examples from a language together with, for every prefix,…
cs.CL2025
Uncertainty-Aware Step-wise Verification with Generative Reward Models
Zihuiwen Ye, Luckeciano Carvalho Melo, Younesse Kaddar +3
Complex multi-step reasoning tasks, such as solving mathematical problems, remain challenging for large language models (LLMs). While outcome supervision is commonly used, process…
cs.CL2024
Improving Reward Models with Synthetic Critiques
Zihuiwen Ye, Fraser Greenlee-Scott, Max Bartolo +3
Reward models (RMs) play a critical role in aligning language models through the process of reinforcement learning from human feedback. RMs are trained to predict a score reflectin…