5 citations · 18 across the 12 of their papers we have counts for
10 papers · 1 filter
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
Aditya Soni, Mayukh Das, Anjaly Parayil +8
The difficulty of exploring and training online on real production systems limits the scope of real-time online data/feedback-driven decision making. The most feasible approach is…
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Corby Rosset, Ching-An Cheng, Arindam Mitra +3
This paper studies post-training large language models (LLMs) using preference feedback from a powerful oracle to help a model iteratively improve over itself. The typical approach…
Survival Instinct in Offline Reinforcement Learning
Anqi Li, Dipendra Misra, Andrey Kolobov +1
We present a novel observation about the behavior of offline reinforcement learning (RL) algorithms: on many benchmark datasets, offline RL can produce well-performing and safe pol…
Improving Offline RL by Blending Heuristics
Sinong Geng, Aldo Pacchiano, Andrey Kolobov +1
We propose Heuristic Blending (HUBL), a simple performance-improving technique for a broad class of offline RL algorithms based on value bootstrapping. HUBL modifies the Bellman op…
MAHALO: Unifying Offline Reinforcement Learning and Imitation Learning from Observations
Anqi Li, Byron Boots, Ching-An Cheng
We study a new paradigm for sequential decision making, called offline policy learning from observations (PLfO). Offline PLfO aims to learn policies using datasets with substandard…
Adversarial Model for Offline Reinforcement Learning
Mohak Bhardwaj, Tengyang Xie, Byron Boots +2
We propose a novel model-based offline Reinforcement Learning (RL) framework, called Adversarial Model for Offline Reinforcement Learning (ARMOR), which can robustly learn policies…