papers

Publications (210)

cs.LG2024

Learning Memory Mechanisms for Decision Making through Demonstrations

William Yue, Bo Liu, Peter Stone

In Partially Observable Markov Decision Processes, integrating an agent's history into memory poses a significant challenge for decision-making. Traditional imitation learning, rel…

cs.AI2025

Automated Reward Design for Gran Turismo

Michel Ma, Takuma Seno, Kaushik Subramanian +3

When designing reinforcement learning (RL) agents, a designer communicates the desired agent behavior through the definition of reward functions - numerical feedback given to the a…

cs.NE2022

Effective Mutation Rate Adaptation through Group Elite Selection

Akarsh Kumar, Bo Liu, Risto Miikkulainen +1

Evolutionary algorithms are sensitive to the mutation rate (MR); no single value of this parameter works well across domains. Self-adaptive MR approaches have been proposed but the…

cs.RO2018

An Architecture for Person-Following using Active Target Search

Minkyu Kim, Miguel Arduengo, Nick Walker +4

This paper addresses a novel architecture for person-following robots using active search. The proposed system can be applied in real-time to general mobile robots for learning fea…

cs.AI2025

RLZero: Direct Policy Inference from Language Without In-Domain Supervision

Harshit Sikchi, Siddhant Agarwal, Pranaya Jajoo +6

The reward hypothesis states that all goals and purposes can be understood as the maximization of a received scalar reward signal. However, in practice, defining such a reward sign…

cs.RO2023

Causal Policy Gradient for Whole-Body Mobile Manipulation

Jiaheng Hu, Peter Stone, Roberto Martín-Martín

Developing the next generation of household robot helpers requires combining locomotion and interaction capabilities, which is generally referred to as mobile manipulation (MoMa).…

cs.LG2021

Firefly Neural Architecture Descent: a General Approach for Growing Neural Networks

Lemeng Wu, Bo Liu, Peter Stone +1

We propose firefly neural architecture descent, a general framework for progressively and dynamically growing neural networks to jointly optimize the networks' parameters and archi…

cs.RO2022

Socially Compliant Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Navigation

Haresh Karnan, Anirudh Nair, Xuesu Xiao +6

Social navigation is the capability of an autonomous agent, such as a robot, to navigate in a 'socially compliant' manner in the presence of other intelligent agents such as humans…

cs.MA2025

Sequence Modeling for N-Agent Ad Hoc Teamwork

Caroline Wang, Di Yang Shi, Elad Liebman +3

N-agent ad hoc teamwork (NAHT) is a newly introduced challenge in multi-agent reinforcement learning, where controlled subteams of varying sizes must dynamically collaborate with v…

cs.AI2018

Behavioral Cloning from Observation

Faraz Torabi, Garrett Warnell, Peter Stone

Humans often learn how to perform tasks via imitation: they observe others perform a task, and then very quickly infer the appropriate actions to take based on their observations.…

cs.LG2023

Event Tables for Efficient Experience Replay

Varun Kompella, Thomas J. Walsh, Samuel Barrett +2

Experience replay (ER) is a crucial component of many deep reinforcement learning (RL) systems. However, uniform sampling from an ER buffer can lead to slow convergence and unstabl…

cs.RO2024

Learning to Look: Seeking Information for Decision Making via Policy Factorization

Shivin Dass, Jiaheng Hu, Ben Abbatematteo +2

Many robot manipulation tasks require active or interactive exploration behavior in order to be performed successfully. Such tasks are ubiquitous in embodied domains, where agents…

cs.AI2012

Empowerment for Continuous Agent-Environment Systems

Tobias Jung, Daniel Polani, Peter Stone

This paper develops generalizations of empowerment to continuous states. Empowerment is a recently introduced information-theoretic quantity motivated by hypotheses about the effic…

cs.RO2025

PACER: Preference-conditioned All-terrain Costmap Generation

Luisa Mao, Garrett Warnell, Peter Stone +1

In autonomous robot navigation, terrain cost assignment is typically performed using a semantics-based paradigm in which terrain is first labeled using a pre-trained semantic class…

cs.AI2024

N-Agent Ad Hoc Teamwork

Caroline Wang, Arrasy Rahman, Ishan Durugkar +2

Current approaches to learning cooperative multi-agent behaviors assume relatively restrictive settings. In standard fully cooperative multi-agent reinforcement learning, the learn…

cs.LG2024

A Super-human Vision-based Reinforcement Learning Agent for Autonomous Racing in Gran Turismo

Miguel Vasco, Takuma Seno, Kenta Kawamoto +3

Racing autonomous cars faster than the best human drivers has been a longstanding grand challenge for the fields of Artificial Intelligence and robotics. Recently, an end-to-end de…

cs.LG2022

Real-world challenges for multi-agent reinforcement learning in grid-interactive buildings

Kingsley Nweye, Bo Liu, Peter Stone +1

Building upon prior research that highlighted the need for standardizing environments for building control research, and inspired by recently introduced challenges for real life re…

cs.AI2017

Data-Efficient Policy Evaluation Through Behavior Policy Search

Josiah P. Hanna, Philip S. Thomas, Peter Stone +1

We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its per…

cs.RO2020

Stochastic Grounded Action Transformation for Robot Learning in Simulation

Siddharth Desai, Haresh Karnan, Josiah P. Hanna +2

Robot control policies learned in simulation do not often transfer well to the real world. Many existing solutions to this sim-to-real problem, such as the Grounded Action Transfor…

cs.CY2019

Teaching Social Behavior through Human Reinforcement for Ad hoc Teamwork -The STAR Framework

Shani Alkoby, Avilash Rath, Peter Stone

As AI technology continues to develop, more and more agents will become capable of long term autonomy alongside people. Thus, a recent line of research has studied the problem of t…

cs.RO2023

STERLING: Self-Supervised Terrain Representation Learning from Unconstrained Robot Experience

Haresh Karnan, Elvin Yang, Daniel Farkash +3

Terrain awareness, i.e., the ability to identify and distinguish different types of terrain, is a critical ability that robots must have to succeed at autonomous off-road navigatio…

cs.LG2025

Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy

Bram Grooten, Patrick MacAlpine, Kaushik Subramanian +2

Generalization to unseen environments is a significant challenge in the field of robotics and control. In this work, we focus on contextual reinforcement learning, where agents act…

cs.AI2020

Learning and Reasoning for Robot Dialog and Navigation Tasks

Keting Lu, Shiqi Zhang, Peter Stone +1

Reinforcement learning and probabilistic reasoning algorithms aim at learning from interaction experiences and reasoning with probabilistic contextual knowledge respectively. In th…

cs.LG2025

Learning a Fast Mixing Exogenous Block MDP using a Single Trajectory

Alexander Levine, Peter Stone, Amy Zhang

In order to train agents that can quickly adapt to new objectives or reward functions, efficient unsupervised representation learning in sequential decision-making environments can…

cs.RO2024

Robot Air Hockey: A Manipulation Testbed for Robot Learning with Reinforcement Learning

Caleb Chuck, Carl Qi, Michael J. Munje +13

Reinforcement Learning is a promising tool for learning complex policies even in fast-moving and object-interactive domains where human teleoperation or hard-coded policies might f…

cs.RO2018

Interaction and Autonomy in RoboCup@Home and Building-Wide Intelligence

Justin Hart, Harel Yedidsion, Yuqian Jiang +8

Efforts are underway at UT Austin to build autonomous robot systems that address the challenges of long-term deployments in office environments and of the more prescribed domestic…

cs.RO2025

Deadlock-free, Safe, and Decentralized Multi-Robot Navigation in Social Mini-Games via Discrete-Time Control Barrier Functions

Rohan Chandra, Vrushabh Zinage, Efstathios Bakolas +2

We present an approach to ensure safe and deadlock-free navigation for decentralized multi-robot systems operating in constrained environments, including doorways and intersections…

cs.AI2012

Gaussian Processes for Sample Efficient Reinforcement Learning with RMAX-like Exploration

Tobias Jung, Peter Stone

We present an implementation of model-based online reinforcement learning (RL) for continuous domains with deterministic transitions that is specifically designed to achieve low sa…

cs.RO2025

ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion

Zichao Hu, Chen Tang, Michael J. Munje +6

This paper considers the problem of enabling robots to navigate dynamic environments while following instructions. The challenge lies in the combinatorial nature of instruction spe…

cs.AI2020

An Imitation from Observation Approach to Transfer Learning with Dynamics Mismatch

Siddharth Desai, Ishan Durugkar, Haresh Karnan +3

We examine the problem of transferring a policy learned in a source environment to a target environment with different dynamics, particularly in the case where it is critical to re…

cs.LG2023

-Policy Gradients: A General Framework for Goal Conditioned RL using -Divergences

Siddhant Agarwal, Ishan Durugkar, Peter Stone +1

Goal-Conditioned Reinforcement Learning (RL) problems often have access to sparse rewards where the agent receives a reward signal only when it has achieved the goal, making policy…

cs.RO2021

APPLE: Adaptive Planner Parameter Learning from Evaluative Feedback

Zizhao Wang, Xuesu Xiao, Garrett Warnell +1

Classical autonomous navigation systems can control robots in a collision-free manner, oftentimes with verifiable safety and explainability. When facing new environments, however,…

cs.CV2022

COOPERNAUT: End-to-End Driving with Cooperative Perception for Networked Vehicles

Jiaxun Cui, Hang Qiu, Dian Chen +2

Optical sensors and learning algorithms for autonomous vehicles have dramatically advanced in the past few years. Nonetheless, the reliability of today's autonomous vehicles is hin…

cs.RO2022

Adversarial Imitation Learning from Video using a State Observer

Haresh Karnan, Garrett Warnell, Faraz Torabi +1

The imitation learning research community has recently made significant progress towards the goal of enabling artificial agents to imitate behaviors from video demonstrations alone…

cs.RO2022

Bottom-Up Skill Discovery from Unsegmented Demonstrations for Long-Horizon Robot Manipulation

Yifeng Zhu, Peter Stone, Yuke Zhu

We tackle real-world long-horizon robot manipulation tasks through skill discovery. We present a bottom-up approach to learning a library of reusable skills from unsegmented demons…

cs.LG2023

Metric Residual Networks for Sample Efficient Goal-Conditioned Reinforcement Learning

Bo Liu, Yihao Feng, Qiang Liu +1

Goal-conditioned reinforcement learning (GCRL) has a wide range of potential real-world applications, including manipulation and navigation problems in robotics. Especially in such…

cs.RO2026

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

Zifan Xu, Ran Gong, Maria Vittoria Minniti +10

Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (e.g., through teleoperation) serve as…

cs.LG2022

Value Function Decomposition for Iterative Design of Reinforcement Learning Agents

James MacGlashan, Evan Archer, Alisa Devlic +4

Designing reinforcement learning (RL) agents is typically a difficult process that requires numerous design iterations. Learning can fail for a multitude of reasons, and standard R…

cs.RO2020

Reinforced Grounded Action Transformation for Sim-to-Real Transfer

Haresh Karnan, Siddharth Desai, Josiah P. Hanna +2

Robots can learn to do complex tasks in simulation, but often, learned behaviors fail to transfer well to the real world due to simulator imperfections (the reality gap). Some exis…

cs.AI2021

Expected Value of Communication for Planning in Ad Hoc Teamwork

William Macke, Reuth Mirsky, Peter Stone

A desirable goal for autonomous agents is to be able to coordinate on the fly with previously unknown teammates. Known as "ad hoc teamwork", enabling such a capability has been rec…

cs.RO2025

Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer

Adam Labiosa, Zhihan Wang, Siddhant Agarwal +10

Robot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) i…

cs.RO2023

Wait, That Feels Familiar: Learning to Extrapolate Human Preferences for Preference Aligned Path Planning

Haresh Karnan, Elvin Yang, Garrett Warnell +2

Autonomous mobility tasks such as lastmile delivery require reasoning about operator indicated preferences over terrains on which the robot should navigate to ensure both robot saf…

cs.LG2026

A Champion-level Vision-based Reinforcement Learning Agent for Competitive Racing in Gran Turismo 7

Hojoon Lee, Takuma Seno, Jun Jet Tai +4

Deep reinforcement learning has achieved superhuman racing performance in high-fidelity simulators like Gran Turismo 7 (GT7). It typically utilizes global features that require ins…

cs.AI2021

Sequential Online Chore Division for Autonomous Vehicle Convoy Formation

Harel Yedidsion, Shani Alkoby, Peter Stone

Chore division is a class of fair division problems in which some undesirable "resource" must be shared among a set of participants, with each participant wanting to get as little…

cs.LG2020

Reinforcement Learning for Optimization of COVID-19 Mitigation policies

Varun Kompella, Roberto Capobianco, Stacy Jong +5

The year 2020 has seen the COVID-19 virus lead to one of the worst global pandemics in history. As a result, governments around the world are faced with the challenge of protecting…

cs.AI2019

Leveraging Human Guidance for Deep Reinforcement Learning Tasks

Ruohan Zhang, Faraz Torabi, Lin Guan +2

Reinforcement learning agents can learn to solve sequential decision tasks by interacting with the environment. Human knowledge of how to solve these tasks can be incorporated usin…

cs.RO2025

PRESTO: Fast Motion Planning Using Diffusion Models Based on Key-Configuration Environment Representation

Mingyo Seo, Yoonyoung Cho, Yoonchang Sung +3

We introduce a learning-guided motion planning framework that generates seed trajectories using a diffusion model for trajectory optimization. Given a workspace, our method approxi…

cs.LG2022

BOME! Bilevel Optimization Made Easy: A Simple First-Order Approach

Mao Ye, Bo Liu, Stephen Wright +2

Bilevel optimization (BO) is useful for solving a variety of important machine learning problems including but not limited to hyperparameter optimization, meta-learning, continual…

cs.LG2026

Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement Learning

Jiaheng Hu, Zizhao Wang, Peter Stone +1

A hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment. However, existing unsupervised skill discovery m…

cs.RO2025

CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning

Jiaxun Cui, Chen Tang, Jarrett Holtz +4

Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication is usually not human-understandable. Usin…

cs.LG2021

Adversarial Intrinsic Motivation for Reinforcement Learning

Ishan Durugkar, Mauricio Tec, Scott Niekum +1

Learning with an objective to minimize the mismatch with a reference distribution has been shown to be useful for generative modeling and imitation learning. In this paper, we inve…

cs.RO2025

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training

Jiaheng Hu, Peter Stone, Roberto Martín-Martín

Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as mobile manipulators. While reinforcement…

cs.RO2023

Motion Planning (In)feasibility Detection using a Prior Roadmap via Path and Cut Search

Yoonchang Sung, Peter Stone

Motion planning seeks a collision-free path in a configuration space (C-space), representing all possible robot configurations in the environment. As it is challenging to construct…

cs.RO2019

Solving Service Robot Tasks: UT Austin Villa@Home 2019 Team Report

Rishi Shah, Yuqian Jiang, Haresh Karnan +8

RoboCup@Home is an international robotics competition based on domestic tasks requiring autonomous capabilities pertaining to a large variety of AI technologies. Research challenge…

cs.AI2026

Coachable agents for interactive gameplay

Roberto Capobianco, Harm van Seijen, Nolan D. Bard +38

Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation m…

cs.LG2026

Factored Latent Action World Models

Zizhao Wang, Chang Shi, Jiaheng Hu +4

Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions provide a natural interface for u…

cs.RO2024

TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation

Shivin Dass, Wensi Ai, Yuqian Jiang +6

A critical bottleneck limiting imitation learning in robotics is the lack of data. This problem is more severe in mobile manipulation, where collecting demonstrations is harder tha…

cs.LG2022

Learning a Shield from Catastrophic Action Effects: Never Repeat the Same Mistake

Shahaf S. Shperberg, Bo Liu, Peter Stone

Agents that operate in an unknown environment are bound to make mistakes while learning, including, at least occasionally, some that lead to catastrophic consequences. When humans…

cs.RO2021

Agile Robot Navigation through Hallucinated Learning and Sober Deployment

Xuesu Xiao, Bo Liu, Peter Stone

Learning from Hallucination (LfH) is a recent machine learning paradigm for autonomous navigation, which uses training data collected in completely safe environments and adds numer…

cs.AI2022

Scalable Multiagent Driving Policies For Reducing Traffic Congestion

Jiaxun Cui, William Macke, Harel Yedidsion +2

Traffic congestion is a major challenge in modern urban settings. The industry-wide development of autonomous and automated vehicles (AVs) motivates the question of how can AVs con…

cs.HC2025

Harmful Traits of AI Companions

W. Bradley Knox, Katie Bradford, Samanta Varela Castro +6

Amid the growing prevalence of human-AI interaction, large language models and other AI-based entities increasingly provide forms of companionship to human users. Such AI companion…

cs.RO2021

APPLI: Adaptive Planner Parameter Learning From Interventions

Zizhao Wang, Xuesu Xiao, Bo Liu +2

While classical autonomous navigation systems can typically move robots from one point to another safely and in a collision-free manner, these systems may fail or produce suboptima…

cs.AI2025

The Essentials of AI for Life and Society: An AI Literacy Course for the University Community

Joydeep Biswas, Don Fussell, Peter Stone +4

We describe the development of a one-credit course to promote AI literacy at The University of Texas at Austin. In response to a call for the rapid deployment of class to serve a b…

cs.LG2020

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Sanmit Narvekar, Bei Peng, Matteo Leonetti +3

Reinforcement learning (RL) is a popular paradigm for addressing sequential decision tasks in which the agent has only limited environmental feedback. Despite many advances over th…

math.OC2020

Policy Evaluation in Continuous MDPs with Efficient Kernelized Gradient Temporal Difference

Alec Koppel, Garrett Warnell, Ethan Stump +2

We consider policy evaluation in infinite-horizon discounted Markov decision problems (MDPs) with infinite spaces. We reformulate this task a compositional stochastic program with…

cs.RO2024

Grounded Curriculum Learning

Linji Wang, Zifan Xu, Peter Stone +1

The high cost of real-world data for robotics Reinforcement Learning (RL) leads to the wide usage of simulators. Despite extensive work on building better dynamics models for simul…

cs.LG2021

Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy

Yunshu Du, Garrett Warnell, Assefaw Gebremedhin +2

Experience replay (ER) improves the data efficiency of off-policy reinforcement learning (RL) algorithms by allowing an agent to store and reuse its past experiences in a replay bu…

cs.RO2022

Autonomous Ground Navigation in Highly Constrained Spaces: Lessons learned from The BARN Challenge at ICRA 2022

Xuesu Xiao, Zifan Xu, Zizhao Wang +14

The BARN (Benchmark Autonomous Robot Navigation) Challenge took place at the 2022 IEEE International Conference on Robotics and Automation (ICRA 2022) in Philadelphia, PA. The aim…

cs.LG2025

Adversarial Reinforcement Learning for Large Language Model Agent Safety

Zizhao Wang, Dingcheng Li, Vaishakh Keshava +4

Large Language Model (LLM) agents can leverage tools such as Google Search to complete complex tasks. However, this tool usage introduces the risk of indirect prompt injections, wh…

cs.AI2024

Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning

Zizhao Wang, Caroline Wang, Xuesu Xiao +2

Two desiderata of reinforcement learning (RL) algorithms are the ability to learn from relatively little experience and the ability to learn policies that generalize to a range of…

cs.RO2025

Multi-Agent Inverse Reinforcement Learning in Real World Unstructured Pedestrian Crowds

Rohan Chandra, Haresh Karnan, Negar Mehr +2

Social robot navigation in crowded public spaces such as university campuses, restaurants, grocery stores, and hospitals, is an increasingly important area of research. One of the…

cs.AI2023

LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Bo Liu, Yifeng Zhu, Chongkai Gao +4

Lifelong learning offers a promising paradigm of building a generalist agent that learns and adapts over its lifespan. Unlike traditional lifelong learning problems in image and te…

cs.LG2024

Fine-Grained Gradient Restriction: A Simple Approach for Mitigating Catastrophic Forgetting

Bo Liu, Mao Ye, Peter Stone +1

A fundamental challenge in continual learning is to balance the trade-off between learning new tasks and remembering the previously acquired knowledge. Gradient Episodic Memory (GE…

cs.RO2019

Desiderata for Planning Systems in General-Purpose Service Robots

Nick Walker, Yuqian Jiang, Maya Cakmak +1

General-purpose service robots are expected to undertake a broad range of tasks at the request of users. Knowledge representation and planning systems are essential to flexible aut…

cs.CL2018

Learning a Policy for Opportunistic Active Learning

Aishwarya Padmakumar, Peter Stone, Raymond J. Mooney

Active learning identifies data points to label that are expected to be the most useful in improving a supervised model. Opportunistic active learning incorporates active learning…

cs.LG2026

Influencing Humans to Conform to Preference Models for RLHF

Stephane Hatgis-Kessell, W. Bradley Knox, Serena Booth +1

Designing a reinforcement learning from human feedback (RLHF) algorithm to approximate a human's unobservable reward function requires assuming, implicitly or explicitly, a model o…

cs.AI2023

Learning a Robust Multiagent Driving Policy for Traffic Congestion Reduction

Yulin Zhang, William Macke, Jiaxun Cui +2

In most modern cities, traffic congestion is one of the most salient societal challenges. Past research has shown that inserting a limited number of autonomous vehicles (AVs) withi…

stat.ML2024

Sample Efficient Myopic Exploration Through Multitask Reinforcement Learning with Diverse Tasks

Ziping Xu, Zifan Xu, Runxuan Jiang +2

Multitask Reinforcement Learning (MTRL) approaches have gained increasing attention for its wide applications in many important Reinforcement Learning (RL) tasks. However, while re…

cs.LG2020

Reducing Sampling Error in Batch Temporal Difference Learning

Brahma Pavse, Ishan Durugkar, Josiah Hanna +1

Temporal difference (TD) learning is one of the main foundations of modern reinforcement learning. This paper studies the use of TD(0), a canonical TD algorithm, to estimate the va…

cs.AI2026

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

Joydeep Biswas, Sheila Schoepp, Gautham Vasan +10

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances i…

cs.LG2025

Dyn-O: Building Structured World Models with Object-Centric Representations

Zizhao Wang, Kaixin Wang, Li Zhao +2

World models aim to capture the dynamics of the environment, enabling agents to predict and plan for future states. In most scenarios of interest, the dynamics are highly centered…

cs.RO2021

A Scavenger Hunt for Service Robots

Harel Yedidsion, Jennifer Suriadinata, Zifan Xu +2

Creating robots that can perform general-purpose service tasks in a human-populated environment has been a longstanding grand challenge for AI and Robotics research. One particular…

cs.LG2018

Learning Curriculum Policies for Reinforcement Learning

Sanmit Narvekar, Peter Stone

Curriculum learning in reinforcement learning is a training methodology that seeks to speed up learning of a difficult target task, by first training on a series of simpler tasks a…

cs.RO2025

MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention

Yuxin Chen, Chen Tang, Jianglan Wei +6

Aligning robot behavior with human preferences is crucial for deploying embodied AI agents in human-centered environments. A promising solution is interactive imitation learning fr…

cs.RO2023

Exploring the Cost of Interruptions in Human-Robot Teaming

Swathi Mannem, William Macke, Peter Stone +1

Productive and efficient human-robot teaming is a highly desirable ability in service robots, yet there is a fundamental trade-off that a robot needs to consider in such tasks. On…

cs.HC2020

The EMPATHIC Framework for Task Learning from Implicit Human Feedback

Yuchen Cui, Qiping Zhang, Alessandro Allievi +3

Reactions such as gestures, facial expressions, and vocalizations are an abundant, naturally occurring channel of information that humans provide during interactions. A robot or ot…

cs.RO2022

Visually Grounded Task and Motion Planning for Mobile Manipulation

Xiaohan Zhang, Yifeng Zhu, Yan Ding +3

Task and motion planning (TAMP) algorithms aim to help robots achieve task-level goals, while maintaining motion-level feasibility. This paper focuses on TAMP domains that involve…

cs.LG2023

Task Phasing: Automated Curriculum Learning from Demonstrations

Vaibhav Bajaj, Guni Sharon, Peter Stone

Applying reinforcement learning (RL) to sparse reward domains is notoriously challenging due to insufficient guiding signals. Common RL techniques for addressing such domains inclu…

cs.AI2023

Utilizing Mood-Inducing Background Music in Human-Robot Interaction

Elad Liebman, Peter Stone

Past research has clearly established that music can affect mood and that mood affects emotional and cognitive processing, and thus decision-making. It follows that if a robot inte…

cs.RO2023

VIOLA: Imitation Learning for Vision-Based Manipulation with Object Proposal Priors

Yifeng Zhu, Abhishek Joshi, Peter Stone +1

We introduce VIOLA, an object-centric imitation learning approach to learning closed-loop visuomotor policies for robot manipulation. Our approach constructs object-centric represe…

cs.RO2023

Benchmarking Reinforcement Learning Techniques for Autonomous Navigation

Zifan Xu, Bo Liu, Xuesu Xiao +2

Deep reinforcement learning (RL) has brought many successes for autonomous robot navigation. However, there still exists important limitations that prevent real-world use of RL-bas…

cs.RO2021

Learning Inverse Kinodynamics for Accurate High-Speed Off-Road Navigation on Unstructured Terrain

Xuesu Xiao, Joydeep Biswas, Peter Stone

This paper presents a learning-based approach to consider the effect of unobservable world states in kinodynamic motion planning in order to enable accurate high-speed off-road nav…

cs.RO2025

Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks

Viraj Joshi, Zifan Xu, Bo Liu +2

Multi-task Reinforcement Learning (MTRL) has emerged as a critical training paradigm for applying reinforcement learning (RL) to a set of complex real-world robotic tasks, which de…

cs.AI2026

VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon

Cameron Angliss, Jiaxun Cui, Jiaheng Hu +2

Developing AI agents that can robustly adapt to varying strategic landscapes without retraining is a central challenge in multi-agent learning. Pokémon Video Game Championships (V…

cs.AI2015

Representative Selection in Non Metric Datasets

Elad Liebman, Benny Chor, Peter Stone

This paper considers the problem of representative selection: choosing a subset of data points from a dataset that best represents its overall set of elements. This subset needs to…

cs.RO2026

Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input

Zifan Xu, Myoungkyu Seo, Dongmyeong Lee +8

Learning fast and robust ball-kicking skills is a critical capability for humanoid soccer robots, yet it remains a challenging problem due to the need for rapid leg swings, postura…

cs.LG2025

SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning

Hojoon Lee, Dongyoon Hwang, Donghu Kim +7

Recent advances in CV and NLP have been largely driven by scaling up the number of network parameters, despite traditional theories suggesting that larger networks are prone to ove…

cs.AI2023

Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks

Yuqian Jiang, Sudarshanan Bharadwaj, Bo Wu +3

In continuing tasks, average-reward reinforcement learning may be a more appropriate problem formulation than the more common discounted reward formulation. As usual, learning an o…

cs.RO2022

Learning Perceptual Hallucination for Multi-Robot Navigation in Narrow Hallways

Jin-Soo Park, Xuesu Xiao, Garrett Warnell +2

While current systems for autonomous robot navigation can produce safe and efficient motion plans in static environments, they usually generate suboptimal behaviors when multiple r…

cs.NE2018

Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity inspired by Network Science

Decebal Constantin Mocanu, Elena Mocanu, Peter Stone +3

Through the success of deep learning in various domains, artificial neural networks are currently among the most used artificial intelligence methods. Taking inspiration from the n…