1 citations · 1 across the 1 of their papers we have counts for
3 papers
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
Cansu Sancaktar, David Zhang, Gabriel Synnaeve +1
Reinforcement learning (RL) has emerged as a powerful paradigm for improving large language models beyond supervised fine-tuning, yet sustaining performance gains at scale remains…
GASP: Guided Asymmetric Self-Play For Coding LLMs
Swadesh Jana, Cansu Sancaktar, Tomáš Daniš +3
Asymmetric self-play has emerged as a promising paradigm for post-training large language models, where a teacher continually generates questions for a student to solve at the edge…
Real Robot Challenge 2022: Learning Dexterous Manipulation from Offline Data in the Real World
Nico Gürtler, Felix Widmaier, Cansu Sancaktar +21
Experimentation on real robots is demanding in terms of time and costs. For this reason, a large part of the reinforcement learning (RL) community uses simulators to develop and be…