activity
20212025
most citedSafe Reinforcement Learning Using Advantage-Based Intervention

5 citations · 18 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2024

Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC

Aditya Soni, Mayukh Das, Anjaly Parayil +8

The difficulty of exploring and training online on real production systems limits the scope of real-time online data/feedback-driven decision making. The most feasible approach is…

cs.LG2024★ 2 cited

Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Corby Rosset, Ching-An Cheng, Arindam Mitra +3

This paper studies post-training large language models (LLMs) using preference feedback from a powerful oracle to help a model iteratively improve over itself. The typical approach…

cs.LG2023★ 1 cited

Survival Instinct in Offline Reinforcement Learning

Anqi Li, Dipendra Misra, Andrey Kolobov +1

We present a novel observation about the behavior of offline reinforcement learning (RL) algorithms: on many benchmark datasets, offline RL can produce well-performing and safe pol…

cs.LG2023★ 1 cited

Improving Offline RL by Blending Heuristics

Sinong Geng, Aldo Pacchiano, Andrey Kolobov +1

We propose Heuristic Blending (HUBL), a simple performance-improving technique for a broad class of offline RL algorithms based on value bootstrapping. HUBL modifies the Bellman op…

cs.LG2023

MAHALO: Unifying Offline Reinforcement Learning and Imitation Learning from Observations

Anqi Li, Byron Boots, Ching-An Cheng

We study a new paradigm for sequential decision making, called offline policy learning from observations (PLfO). Offline PLfO aims to learn policies using datasets with substandard…

cs.LG2023★ 4 cited

Adversarial Model for Offline Reinforcement Learning

Mohak Bhardwaj, Tengyang Xie, Byron Boots +2

We propose a novel model-based offline Reinforcement Learning (RL) framework, called Adversarial Model for Offline Reinforcement Learning (ARMOR), which can robustly learn policies…