activity
20162023
most citedDecision Transformer: Reinforcement Learning via Sequence Modeling

465 citations · 827 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

18 papers · 1 filter

cs.LG2023

Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models

Hritik Bansal, John Dang, Aditya Grover

Aligning large language models (LLMs) with human values and intents critically involves the use of human or AI feedback. While dense feedback annotations are expensive to acquire a…

cs.LG2023

ClimateLearn: Benchmarking Machine Learning for Weather and Climate Modeling

Tung Nguyen, Jason Jewik, Hritik Bansal +2

Modeling weather and climate is an essential endeavor to understand the near- and long-term impacts of climate change, as well as inform technology and policymaking for adaptation…

cs.LG2023

Diffusion Models for Black-Box Optimization

Siddarth Krishnamoorthy, Satvik Mehul Mashkaria, Aditya Grover

The goal of offline black-box optimization (BBO) is to optimize an expensive black-box function using a fixed dataset of function evaluations. Prior works consider forward approach…

cs.LG2023

Decision Stacks: Flexible Reinforcement Learning via Modular Generative Models

Siyan Zhao, Aditya Grover

Reinforcement learning presents an attractive paradigm to reason about several distinct aspects of sequential decision making, such as specifying complex goals, planning future obs…

cs.LG20231 cited

Scaling Pareto-Efficient Decision Making Via Offline Multi-Objective RL

Baiting Zhu, Meihua Dang, Aditya Grover

The goal of multi-objective reinforcement learning (MORL) is to learn policies that simultaneously optimize multiple competing objectives. In practice, an agent's preferences over…

cs.LG20222 cited

Imitating, Fast and Slow: Robust learning from demonstrations via decision-time planning

Carl Qi, Pieter Abbeel, Aditya Grover

The goal of imitation learning is to mimic expert behavior from demonstrations, without access to an explicit reward signal. A popular class of approach infers the (unknown) reward…