5 papers
Offline Constrained RLHF with Multiple Preference Oracles
Brenden Latham, Mehrdad Moharrami
We study offline constrained reinforcement learning from human feedback with multiple preference oracles. Motivated by applications that trade off performance with safety or fairne…
Tail Distribution of Regret in Optimistic Reinforcement Learning
Sajad Khodadadian, Mehrdad Moharrami
We derive instance-dependent tail bounds for the regret of optimism-based reinforcement learning in finite-horizon tabular Markov decision processes with unknown transition dynamic…
Learning to Admit Optimally in an Queueing System with Unknown Service Rate
Saghar Adler, Mehrdad Moharrami, Vijay Subramanian
Motivated by applications of the Erlang-B blocking model and the extended model that allows for some queueing, beyond communication networks to sizing and pricing in pr…
The Planted Spanning Tree Problem
Mehrdad Moharrami, Cristopher Moore, Jiaming Xu
We study the problem of detecting and recovering a planted spanning tree hidden within a complete, randomly weighted graph . Specifically, each edge has a non-nega…
A Stackelberg Game Model of Flocking
Chenlan Wang, Mehrdad Moharrami, Mingyan Liu
We study a Stackelberg game to examine how two agents determine to cooperate while competing with each other. Each selects an arrival time to a destination, the earlier one fetchin…