1 citations · 1 across the 4 of their papers we have counts for
7 papers · 1 filter
GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance
Jiri Navratil, Jarret Ross, Payel Das +4
The ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology…
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
Youssef Mroueh
Group Relative Policy Optimization (GRPO) was introduced and used recently for promoting reasoning in LLMs under verifiable (binary) rewards. We show that the mean + variance calib…
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
Gholamali Aminian, Amir R. Asadi, Idan Shenfeld +1
Recent methods for aligning large language models (LLMs) with human feedback predominantly rely on a single reference model, which limits diversity, model overfitting, and underuti…
Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
Nicolas Dupuis, Adarsh Tiwari, Youssef Mroueh +3
Qiskit is an open-source quantum computing framework that allows users to design, simulate, and run quantum circuits on real quantum hardware. We explore post-training techniques f…
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
Youssef Mroueh, Nicolas Dupuis, Brian Belgodere +6
We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Opti…
Gradient Flows and Riemannian Structure in the Gromov-Wasserstein Geometry
Zhengxin Zhang, Ziv Goldfeld, Kristjan Greenewald +2
The Wasserstein space of probability measures is known for its intricate Riemannian structure, which underpins the Wasserstein geometry and enables gradient flow algorithms. Howeve…