activity
20192026
collaborators

13 papers

cs.AI2026

Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution

Han Li, Jinyu Tian, Rili Feng +10

Large language models (LLMs) still struggle with the rigorous reasoning demands of hard competitive programming. While recent multi-agent frameworks attempt to bridge this reliabil…

cs.LG2026

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning

Aravind Venugopal, Jiayu Chen, Xudong Wu +3

The temporal lag between actions and their long-term consequences makes credit assignment a challenge when learning goal-directed behaviors from data. Generative world models captu…

cs.LG2026

UpSkill: Mutual Information Skill Learning for Structured Response Diversity in LLMs

Devan Shah, Owen Yang, Daniel Yang +2

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of large language models (LLMs) on mathematics and programming tasks, but standard approa…

cs.LG2026

Can We Really Learn One Representation to Optimize All Rewards?

Chongyi Zheng, Royina Karegoudra Jayanth, Benjamin Eysenbach

As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough theoretical understanding of these methods becomes of equal importance to the…

cs.LG2025

Consistent Zero-Shot Imitation with Contrastive Goal Inference

Kathryn Wantlin, Chongyi Zheng, Benjamin Eysenbach

Zero-shot imitation learning requires an agent to reproduce expert behavior from a single demonstration without additional environment interaction or gradient updates at test time.…

cs.LG2025

Value Flows

Perry Dong, Chongyi Zheng, Chelsea Finn +2

While most reinforcement learning methods today flatten the distribution of future returns to a single scalar value, distributional RL methods exploit the return distribution to pr…