activity
20132026
most citedConvex Learning of Multiple Tasks and their Structure

37 citations · 141 across the 29 of their papers we have counts for

collaborators
Showing 2025Show all

6 papers · 1 filter

quant-ph20251 cited

Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant

Nicolas Dupuis, Adarsh Tiwari, Youssef Mroueh +3

Qiskit is an open-source quantum computing framework that allows users to design, simulate, and run quantum circuits on real quantum hardware. We explore post-training techniques f…

cs.LG2025

GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance

Jiri Navratil, Jarret Ross, Payel Das +4

The ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology…

cs.LG2025

Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis

We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of- test-time scaling with…

cs.LG2025

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

Youssef Mroueh, Nicolas Dupuis, Brian Belgodere +6

We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Opti…

cs.LG2025

Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification

Youssef Mroueh

Group Relative Policy Optimization (GRPO) was introduced and used recently for promoting reasoning in LLMs under verifiable (binary) rewards. We show that the mean + variance calib…

cs.LG2025

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

Gholamali Aminian, Amir R. Asadi, Idan Shenfeld +1

Recent methods for aligning large language models (LLMs) with human feedback predominantly rely on a single reference model, which limits diversity, model overfitting, and underuti…