Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
Khiem Le, Phuc Nguyen, Youssef Mroueh +4
Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models, but it suffers from two critic…
cs.LG2026
CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery
Youssef Mroueh, Carlos Fonseca, Brian Belgodere +1
Scientific algorithm discovery is iterative: hypotheses are proposed, implemented, stress-tested, and revised. Current LLM-guided search systems accelerate proposal generation, but…