4 papers · 1 filter
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards
Saurabh Dash, Pierre Clavier, John Dang +4
Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be checked automatically. However,…
Aya Vision: Advancing the Frontier of Multilingual Multimodality
Saurabh Dash, Yiyang Nan, John Dang +22
Building multimodal language models is fundamentally challenging: it requires aligning vision and language modalities, curating high-quality instruction data, and avoiding the degr…
Improving Reward Models with Synthetic Critiques
Zihuiwen Ye, Fraser Greenlee-Scott, Max Bartolo +3
Reward models (RMs) play a critical role in aligning language models through the process of reinforcement learning from human feedback. RMs are trained to predict a score reflectin…
LLMCRIT: Teaching Large Language Models to Use Criteria
Weizhe Yuan, Pengfei Liu, Matthias Gallé
Humans follow criteria when they execute tasks, and these criteria are directly used to assess the quality of task completion. Therefore, having models learn to use criteria to pro…