6 papers · 1 filter
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
Utsav Singh, Souradip Chakraborty, Wesley A. Suttle +6
Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundament…
Deceptive Exploration in Multi-armed Bandits
I. Arda Vurankaya, Mustafa O. Karabag, Wesley A. Suttle +3
We consider a multi-armed bandit setting in which each arm has a public and a private reward distribution. An observer expects an agent to follow Thompson Sampling according to the…
Signal attenuation enables scalable decentralized multi-agent reinforcement learning over networks
Wesley A Suttle, Vipul K Sharma, Brian M Sadler
Multi-agent reinforcement learning (MARL) methods typically require that agents enjoy global state observability, preventing development of decentralized algorithms and limiting sc…
Value of Information-based Deceptive Path Planning Under Adversarial Interventions
Wesley A. Suttle, Jesse Milzman, Mustafa O. Karabag +2
Existing methods for deceptive path planning (DPP) address the problem of designing paths that conceal their true goal from a passive, external observer. Such methods do not apply…
Behavioral Entropy-Guided Dataset Generation for Offline Reinforcement Learning
Wesley A. Suttle, Aamodh Suresh, Carlos Nieto-Granda
Entropy-based objectives are widely used to perform state space exploration in reinforcement learning (RL) and dataset generation for offline RL. Behavioral entropy (BE), a rigorou…
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
Utsav Singh, Souradip Chakraborty, Wesley A. Suttle +3
Learning control policies to perform complex robotics tasks from human preference data presents significant challenges. On the one hand, the complexity of such tasks typically requ…