37 citations · 141 across the 29 of their papers we have counts for
6 papers · 1 filter
Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
Nicolas Dupuis, Adarsh Tiwari, Youssef Mroueh +3
Qiskit is an open-source quantum computing framework that allows users to design, simulate, and run quantum circuits on real quantum hardware. We explore post-training techniques f…
GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance
Jiri Navratil, Jarret Ross, Payel Das +4
The ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology…
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis
We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of- test-time scaling with…
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
Youssef Mroueh, Nicolas Dupuis, Brian Belgodere +6
We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Opti…
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
Youssef Mroueh
Group Relative Policy Optimization (GRPO) was introduced and used recently for promoting reasoning in LLMs under verifiable (binary) rewards. We show that the mean + variance calib…
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
Gholamali Aminian, Amir R. Asadi, Idan Shenfeld +1
Recent methods for aligning large language models (LLMs) with human feedback predominantly rely on a single reference model, which limits diversity, model overfitting, and underuti…