1 citations · 2 across the 8 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Entropy-Preserving Reinforcement Learning
Aleksei Petrenko, Ben Lipkin, Kevin Chen +4
Policy gradient algorithms have driven many recent advancements in language model reasoning. An appealing property is their ability to learn from exploration on their own trajector…
cs.LG2025
Reinforcement Learning for Long-Horizon Interactive LLM Agents
Kevin Chen, Marco Cusumano-Towner, Brody Huval +4
Interactive digital agents (IDAs) leverage APIs of stateful digital environments to perform tasks in response to user requests. While IDAs powered by instruction-tuned large langua…