4 papers
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards
Saurabh Dash, Pierre Clavier, John Dang +4
Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be checked automatically. However,…
EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
Daniel Scalena, Leonidas Zotos, Elisabetta Fersini +2
With the rise of reasoning language models and test-time scaling methods as a paradigm for improving model performance, substantial computation is often required to generate multip…
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
Yijiang River Dong, Tiancheng Hu, Yinhong Liu +2
While Reinforcement Learning from Human Feedback (RLHF) is widely used to align Large Language Models (LLMs) with human preferences, it typically assumes homogeneous preferences ac…
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
Abdullatif Köksal, Marion Thaler, Ayyoob Imani +3
Instruction tuning enhances large language models (LLMs) by aligning them with human preferences across diverse tasks. Traditional approaches to create instruction tuning datasets…