activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach

Utsav Singh, Souradip Chakraborty, Wesley A. Suttle +6

Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundament…

cs.LG2025

Deceptive Exploration in Multi-armed Bandits

I. Arda Vurankaya, Mustafa O. Karabag, Wesley A. Suttle +3

We consider a multi-armed bandit setting in which each arm has a public and a private reward distribution. An observer expects an agent to follow Thompson Sampling according to the…

cs.LG2025

Signal attenuation enables scalable decentralized multi-agent reinforcement learning over networks

Wesley A Suttle, Vipul K Sharma, Brian M Sadler

Multi-agent reinforcement learning (MARL) methods typically require that agents enjoy global state observability, preventing development of decentralized algorithms and limiting sc…

cs.LG2025

Value of Information-based Deceptive Path Planning Under Adversarial Interventions

Wesley A. Suttle, Jesse Milzman, Mustafa O. Karabag +2

Existing methods for deceptive path planning (DPP) address the problem of designing paths that conceal their true goal from a passive, external observer. Such methods do not apply…

cs.LG2025

Behavioral Entropy-Guided Dataset Generation for Offline Reinforcement Learning

Wesley A. Suttle, Aamodh Suresh, Carlos Nieto-Granda

Entropy-based objectives are widely used to perform state space exploration in reinforcement learning (RL) and dataset generation for offline RL. Behavioral entropy (BE), a rigorou…

cs.LG2024

DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning

Utsav Singh, Souradip Chakraborty, Wesley A. Suttle +3

Learning control policies to perform complex robotics tasks from human preference data presents significant challenges. On the one hand, the complexity of such tasks typically requ…