3 papers
cs.LG2026
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
Khiem Le, Phuc Nguyen, Youssef Mroueh +4
Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models, but it suffers from two critic…
cs.LG2026
CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery
Youssef Mroueh, Carlos Fonseca, Brian Belgodere +1
Scientific algorithm discovery is iterative: hypotheses are proposed, implemented, stress-tested, and revised. Current LLM-guided search systems accelerate proposal generation, but…
stat.ML2025
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
Gholamali Aminian, Idan Shenfeld, Amir R. Asadi +2
A simple yet effective method for inference-time alignment of generative models is Best-of- (BoN), where outcomes are sampled from a reference policy, evaluated using a prox…