activity
20172026
most citedExtrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations

30 citations · 93 across the 34 of their papers we have counts for

collaborators
Showing cs.AIShow all

10 papers · 1 filter

cs.AI2026

Hierarchical Experimentalist Agents

Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka +2

Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-tra…

cs.AI2026

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal +3

Behavioral Foundation Models (BFMs) produce agents with the capability to adapt to any unknown reward or task. These methods, however, are only able to produce near-optimal policie…

cs.AI2025

Evaluation-Aware Reinforcement Learning

Shripad Vilasrao Deshmukh, Will Schwarzer, Scott Niekum

Policy evaluation is a core component of many reinforcement learning (RL) algorithms and a critical tool for ensuring safe deployment of RL policies. However, existing policy evalu…

cs.AI2025

A Descriptive and Normative Theory of Human Beliefs in RLHF

Sylee Dandekar, Shripad Deshmukh, Frank Chiu +2

Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values. In this work, we propose that human belie…

cs.AI2024

RLZero: Direct Policy Inference from Language Without In-Domain Supervision

Harshit Sikchi, Siddhant Agarwal, Pranaya Jajoo +6

The reward hypothesis states that all goals and purposes can be understood as the maximization of a received scalar reward signal. However, in practice, defining such a reward sign…

cs.AI2024

Predicting Future Actions of Reinforcement Learning Agents

Stephen Chung, Scott Niekum, David Krueger

As reinforcement learning agents become increasingly deployed in real-world scenarios, predicting future agent actions and events during deployment is important for facilitating be…