3 papers
cs.AI2025
Toward Agents That Reason About Their Computation
Adrian Orenstein, Jessica Chen, Gwyneth Anne Delos Santos +2
While reinforcement learning agents can achieve superhuman performance in many complex tasks, they typically do not become more computationally efficient as they improve. In contra…
cs.AI2025
Generalization in Monitored Markov Decision Processes (Mon-MDPs)
Montaser Mohammedalamen, Michael Bowling
Reinforcement learning (RL) typically models the interaction between the agent and environment as a Markov decision process (MDP), where the rewards that guide the agent's behavior…
cs.CL2025
KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation
Jiabin Fan, Guoqing Luo, Michael Bowling +1
We propose a novel k-step return estimation method (called KETCHUP) for Reinforcement Learning(RL)-based knowledge distillation (KD) in text generation tasks. Our idea is to induce…