13 papers
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
Han Li, Jinyu Tian, Rili Feng +10
Large language models (LLMs) still struggle with the rigorous reasoning demands of hard competitive programming. While recent multi-agent frameworks attempt to bridge this reliabil…
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning
Aravind Venugopal, Jiayu Chen, Xudong Wu +3
The temporal lag between actions and their long-term consequences makes credit assignment a challenge when learning goal-directed behaviors from data. Generative world models captu…
UpSkill: Mutual Information Skill Learning for Structured Response Diversity in LLMs
Devan Shah, Owen Yang, Daniel Yang +2
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of large language models (LLMs) on mathematics and programming tasks, but standard approa…
Can We Really Learn One Representation to Optimize All Rewards?
Chongyi Zheng, Royina Karegoudra Jayanth, Benjamin Eysenbach
As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough theoretical understanding of these methods becomes of equal importance to the…
Consistent Zero-Shot Imitation with Contrastive Goal Inference
Kathryn Wantlin, Chongyi Zheng, Benjamin Eysenbach
Zero-shot imitation learning requires an agent to reproduce expert behavior from a single demonstration without additional environment interaction or gradient updates at test time.…
Value Flows
Perry Dong, Chongyi Zheng, Chelsea Finn +2
While most reinforcement learning methods today flatten the distribution of future returns to a single scalar value, distributional RL methods exploit the return distribution to pr…