3 papers
cs.LG2025
Reevaluating Policy Gradient Methods for Imperfect-Information Games
Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour +6
In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed nu…
cs.AI2024
RLZero: Direct Policy Inference from Language Without In-Domain Supervision
Harshit Sikchi, Siddhant Agarwal, Pranaya Jajoo +6
The reward hypothesis states that all goals and purposes can be understood as the maximization of a received scalar reward signal. However, in practice, defining such a reward sign…
cs.RO2024
Robot Air Hockey: A Manipulation Testbed for Robot Learning with Reinforcement Learning
Caleb Chuck, Carl Qi, Michael J. Munje +13
Reinforcement Learning is a promising tool for learning complex policies even in fast-moving and object-interactive domains where human teleoperation or hard-coded policies might f…