papers

Publications (70)

cs.AI2015

Sequential Feature Explanations for Anomaly Detection

Md Amran Siddiqui, Alan Fern, Thomas G. Dietterich +1

In many applications, an anomaly detection system presents the most anomalous data instance to a human analyst, who then must determine whether the instance is truly of interest (e…

cs.LG2025

Hierarchical Object-Oriented POMDP Planning for Object Rearrangement

Rajesh Mangannavar, Alan Fern, Prasad Tadepalli

We present an online planning framework and a new benchmark dataset for solving multi-object rearrangement problems in partially observable, multi-room environments. Current object…

cs.HC2019

Explaining Reinforcement Learning to Mere Mortals: An Empirical Study

Andrew Anderson, Jonathan Dodge, Amrita Sadarangani +6

We present a user study to investigate the impact of explanations on non-experts' understanding of reinforcement learning (RL) agents. We investigate both a common RL visualization…

cs.RO2021

Learning Task Space Actions for Bipedal Locomotion

Helei Duan, Jeremy Dao, Kevin Green +3

Recent work has demonstrated the success of reinforcement learning (RL) for training bipedal locomotion policies for real robots. This prior work, however, has focused on learning…

cs.LG2020

Optimizing Discrete Spaces via Expensive Evaluations: A Learning to Search Framework

Aryan Deshwal, Syrine Belakaria, Janardhan Rao Doppa +1

We consider the problem of optimizing expensive black-box functions over discrete spaces (e.g., sets, sequences, graphs). The key challenge is to select a sequence of combinatorial…

cs.LG2020

Dream and Search to Control: Latent Space Planning for Continuous Control

Anurag Koul, Varun V. Kumar, Alan Fern +1

Learning and planning with latent space dynamics has been shown to be useful for sample efficiency in model-based reinforcement learning (MBRL) for discrete and continuous control…

cs.RO2022

Optimizing Bipedal Maneuvers of Single Rigid-Body Models for Reinforcement Learning

Ryan Batke, Fangzhou Yu, Jeremy Dao +4

In this work, we propose a method to generate reduced-order model reference trajectories for general classes of highly dynamic maneuvers for bipedal robots for use in sim-to-real r…

cs.LG2014

Coactive Learning for Locally Optimal Problem Solving

Robby Goetschalckx, Alan Fern, Prasad Tadepalli

Coactive learning is an online problem solving setting where the solutions provided by a solver are interactively improved by a domain expert, which in turn drives learning. In thi…

cs.DB2020

Usable & Scalable Learning Over Relational Data With Automatic Language Bias

Jose Picado, Arash Termehchy, Sudhanshu Pathak +3

Relational databases are valuable resources for learning novel and interesting relations and concepts. In order to constraint the search through the large space of candidate defini…

cs.LG2012

Batch Active Learning via Coordinated Matching

Javad Azimi, Alan Fern, Xiaoli Zhang-Fern +2

Most prior work on active learning of classifiers has focused on sequentially selecting one unlabeled example at a time to be labeled in order to reduce the overall labeling effort…

cs.RO2025

Generating Physically Realistic and Directable Human Motions from Multi-Modal Inputs

Aayam Shrestha, Pan Liu, German Ros +2

This work focuses on generating realistic, physically-based human behaviors from multi-modal inputs, which may only partially specify the desired motion. For example, the input may…

cs.RO2020

Learning Memory-Based Control for Human-Scale Bipedal Locomotion

Jonah Siekmann, Srikar Valluri, Jeremy Dao +4

Controlling a non-statically stable biped is a difficult problem largely due to the complex hybrid dynamics involved. Recent work has demonstrated the effectiveness of reinforcemen…

cs.RO2022

Sim-to-Real Learning of Footstep-Constrained Bipedal Dynamic Walking

Helei Duan, Ashish Malik, Jeremy Dao +5

Recently, work on reinforcement learning (RL) for bipedal robots has successfully learned controllers for a variety of dynamic gaits with robust sim-to-real demonstrations. In orde…

cs.RO2026

No More Marching: Learning Humanoid Locomotion for Short-Range SE(2) Targets

Pranay Dugar, Mohitvishnu S. Gadde, Jonah Siekmann +3

Humanoids operating in real-world workspaces must frequently execute task-driven, short-range movements to SE(2) target poses. To be practical, these transitions must be fast, robu…

cs.LG2023

Attention-based Models for Snow-Water Equivalent Prediction

Krishu K. Thapa, Bhupinderjeet Singh, Supriya Savalkar +3

Snow Water-Equivalent (SWE) -- the amount of water available if snowpack is melted -- is a key decision variable used by water management agencies to make irrigation, flood control…

cs.LG2021

Deep Convolution for Irregularly Sampled Temporal Point Clouds

Erich Merrill, Stefan Lee, Li Fuxin +2

We consider the problem of modeling the dynamics of continuous spatial-temporal processes represented by irregular samples through both space and time. Such processes occur in sens…

cs.RO2026

Multi-Quadruped Cooperative Object Transport: Learning Decentralized Pinch-Lift-Move

Bikram Pandit, Aayam Kumar Shrestha, Alan Fern

We study decentralized cooperative transport using teams of N-quadruped robots with arm that must pinch, lift, and move ungraspable objects through physical contact alone. Unlike p…

cs.RO2026

Humanoid Hanoi: Investigating Shared Whole-Body Control for Skill-Based Box Rearrangement

Minku Kim, Kuan-Chia Chen, Aayam Shrestha +3

We investigate a skill-based framework for humanoid box rearrangement that enables long-horizon execution by sequencing reusable skills at the task level. In our architecture, all…

cs.LG2025

Transfer Learning via Auxiliary Labels with Application to Cold-Hardiness Prediction

Kristen Goebel, Paola Pesantez-Cabrera, Markus Keller +1

Cold temperatures can cause significant frost damage to fruit crops depending on their resilience, or cold hardiness, which changes throughout the dormancy season. This has led to…

cs.AI2022

Beyond Value: CHECKLIST for Testing Inferences in Planning-Based RL

Kin-Ho Lam, Delyar Tabatabai, Jed Irvine +4

Reinforcement learning (RL) agents are commonly evaluated via their expected value over a distribution of test scenarios. Unfortunately, this evaluation approach provides limited e…

cs.AI2021

Contrastive Explanations for Reinforcement Learning via Embedded Self Predictions

Zhengxian Lin, Kim-Ho Lam, Alan Fern

We investigate a deep reinforcement learning (RL) architecture that supports explaining why a learned agent prefers one action over another. The key idea is to learn action-values…

cs.RO2024

Learning Vision-Based Bipedal Locomotion for Challenging Terrain

Helei Duan, Bikram Pandit, Mohitvishnu S. Gadde +4

Reinforcement learning (RL) for bipedal locomotion has recently demonstrated robust gaits over moderate terrains using only proprioceptive sensing. However, such blind controllers…

cs.LG2021

Piecewise-constant Neural ODEs

Sam Greydanus, Stefan Lee, Alan Fern

Neural networks are a popular tool for modeling sequential data but they generally do not treat time as a continuous variable. Neural ODEs represent an important exception: they pa…

cs.LG2018

Open Category Detection with PAC Guarantees

Si Liu, Risheek Garrepalli, Thomas G. Dietterich +2

Open category detection is the problem of detecting "alien" test instances that belong to categories or classes that were not present in the training data. In many applications, re…

cs.AI2019

The Choice Function Framework for Online Policy Improvement

Murugeswari Issakkimuthu, Alan Fern, Prasad Tadepalli

There are notable examples of online search improving over hand-coded or learned policies (e.g. AlphaZero) for sequential decision making. It is not clear, however, whether or not…

cs.LG2025

Graph Neural Network Based Action Ranking for Planning

Rajesh Mangannavar, Stefan Lee, Alan Fern +1

We propose a novel approach to learn relational policies for classical planning based on learning to rank actions. We introduce a new graph representation that explicitly captures…

cs.LG2022

Offline Policy Comparison with Confidence: Benchmarks and Baselines

Anurag Koul, Mariano Phielipp, Alan Fern

Decision makers often wish to use offline historical data to compare sequential-action policies at various world states. Importantly, computational tools should produce confidence…

cs.RO2022

Sim-to-Real Learning for Bipedal Locomotion Under Unsensed Dynamic Loads

Jeremy Dao, Kevin Green, Helei Duan +2

Recent work on sim-to-real learning for bipedal locomotion has demonstrated new levels of robustness and agility over a variety of terrains. However, that work, and most prior bipe…

cs.CV2016

Approximate Policy Iteration for Budgeted Semantic Video Segmentation

Behrooz Mahasseni, Sinisa Todorovic, Alan Fern

This paper formulates and presents a solution to the new problem of budgeted semantic video segmentation. Given a video, the goal is to accurately assign a semantic class label to…

eess.SY2014

A Policy Switching Approach to Consolidating Load Shedding and Islanding Protection Schemes

Rich Meier, Eduardo Cotilla-Sanchez, Alan Fern

In recent years there have been many improvements in the reliability of critical infrastructure systems. Despite these improvements, the power systems industry has seen relatively…

cs.AI2018

Visualizing and Understanding Atari Agents

Sam Greydanus, Anurag Koul, Jonathan Dodge +1

While deep reinforcement learning (deep RL) agents are effective at maximizing rewards, it is often unclear what strategies they use to do so. In this paper, we take a step toward…

cs.CV2021

One Explanation is Not Enough: Structured Attention Graphs for Image Classification

Vivswan Shitole, Li Fuxin, Minsuk Kahng +2

Attention maps are a popular way of explaining the decisions of convolutional networks for image classification. Typically, for each image of interest, a single attention map is pr…

cs.LG2025

Budgeted Online Active Learning with Expert Advice and Episodic Priors

Kristen Goebel, William Solow, Paola Pesantez-Cabrera +2

This paper introduces a novel approach to budgeted online active learning from finite-horizon data streams with extremely limited labeling budgets. In agricultural applications, su…

cs.RO2026

No More Blind Spots: Learning Vision-Based Omnidirectional Bipedal Locomotion for Challenging Terrain

Mohitvishnu S. Gadde, Pranay Dugar, Ashish Malik +1

Effective bipedal locomotion in dynamic environments, such as cluttered indoor spaces or uneven terrain, requires agile and adaptive movement in all directions. This necessitates o…

cs.RO2024

Revisiting Reward Design and Evaluation for Robust Humanoid Standing and Walking

Bart van Marum, Aayam Shrestha, Helei Duan +3

A necessary capability for humanoid robots is the ability to stand and walk while rejecting natural disturbances. Recent progress has been made using sim-to-real reinforcement lear…

cs.RO2021

Sim-to-Real Learning of All Common Bipedal Gaits via Periodic Reward Composition

Jonah Siekmann, Yesh Godse, Alan Fern +1

We study the problem of realizing the full spectrum of bipedal locomotion on a real robot with sim-to-real reinforcement learning (RL). A key challenge of learning legged locomotio…

cs.LG2025

Online Optimization for Offline Safe Reinforcement Learning

Yassine Chemingui, Aryan Deshwal, Alan Fern +2

We study the problem of Offline Safe Reinforcement Learning (OSRL), where the goal is to learn a reward-maximizing policy from fixed data under a cumulative cost constraint. We pro…

cs.RO2025

Evaluating Robots Like Human Infants: A Case Study of Learned Bipedal Locomotion

Devin Crowley, Whitney G. Cole, Christina M. Hospodar +3

Typically, learned robot controllers are trained via relatively unsystematic regimens and evaluated with coarse-grained outcome measures such as average cumulative reward. The typi…

cs.RO2025

Optimizing Bipedal Locomotion for The 100m Dash With Comparison to Human Running

Devin Crowley, Jeremy Dao, Helei Duan +3

In this paper, we explore the space of running gaits for the bipedal robot Cassie. Our first contribution is to present an approach for optimizing gait efficiency across a spectrum…

cs.LG2022

Out-of-Distribution Dynamics Detection: RL-Relevant Benchmarks and Results

Mohamad H Danesh, Alan Fern

We study the problem of out-of-distribution dynamics (OODD) detection, which involves detecting when the dynamics of a temporal process change compared to the training-distribution…

cs.RO2022

Learning Dynamic Bipedal Walking Across Stepping Stones

Helei Duan, Ashish Malik, Mohitvishnu S. Gadde +3

In this work, we propose a learning approach for 3D dynamic bipedal walking when footsteps are constrained to stepping stones. While recent work has shown progress on this problem,…

cs.LG2025

Self-attention-based Diffusion Model for Time-series Imputation in Partial Blackout Scenarios

Mohammad Rafid Ul Islam, Prasad Tadepalli, Alan Fern

Missing values in multivariate time series data can harm machine learning performance and introduce bias. These gaps arise from sensor malfunctions, blackouts, and human error and…

cs.LG2012

Active Imitation Learning via Reduction to I.I.D. Active Learning

Kshitij Judah, Alan Fern, Thomas G. Dietterich

In standard passive imitation learning, the goal is to learn a target policy by passively observing full execution trajectories of it. Unfortunately, generating such trajectories c…

cs.LG2025

Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning

Yassine Chemingui, Aryan Deshwal, Honghao Wei +2

Offline safe reinforcement learning (OSRL) involves learning a decision-making policy to maximize rewards from a fixed batch of training data to satisfy pre-defined safety constrai…

cs.LG2023

Grape Cold Hardiness Prediction via Multi-Task Learning

Aseem Saxena, Paola Pesantez-Cabrera, Rohan Ballapragada +3

Cold temperatures during fall and spring have the potential to cause frost damage to grapevines and other fruit plants, which can significantly decrease harvest yields. To help pre…

cs.AI2021

Identifying Reasoning Flaws in Planning-Based RL Using Tree Explanations

Kin-Ho Lam, Zhengxian Lin, Jed Irvine +5

Enabling humans to identify potential flaws in an agent's decision making is an important Explainable AI application. We consider identifying such flaws in a planning-based deep re…

cs.LG2021

Re-understanding Finite-State Representations of Recurrent Policy Networks

Mohamad H. Danesh, Anurag Koul, Alan Fern +1

We introduce an approach for understanding control policies represented as recurrent neural networks. Recent work has approached this problem by transforming such recurrent policy…

cs.AI2026

A Hybrid Modeling Framework for Crop Prediction Tasks via Dynamic Parameter Calibration and Multi-Task Learning

William Solow, Paola Pesantez-Cabrera, Markus Keller +3

Accurate prediction of crop states (e.g., phenology stages and cold hardiness) is essential for timely farm management decisions such as irrigation, fertilization, and canopy manag…

cs.AI2012

Inferring Strategies from Limited Reconnaissance in Real-time Strategy Games

Jesse Hostetler, Ethan W. Dereszynski, Thomas G. Dietterich +1

In typical real-time strategy (RTS) games, enemy units are visible only when they are within sight range of a friendly unit. Knowledge of an opponent's disposition is limited to wh…

cs.LG2018

Learning Finite State Representations of Recurrent Policy Networks

Anurag Koul, Sam Greydanus, Alan Fern

Recurrent neural networks (RNNs) are an effective representation of control policies for a wide range of reinforcement and imitation learning problems. RNN policies, however, are p…

cs.RO2021

Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning

Jonah Siekmann, Kevin Green, John Warila +2

Accurate and precise terrain estimation is a difficult problem for robot locomotion in real-world environments. Thus, it is useful to have systems that do not depend on accurate es…

cs.CV2021

From Heatmaps to Structural Explanations of Image Classifiers

Li Fuxin, Zhongang Qi, Saeed Khorram +4

This paper summarizes our endeavors in the past few years in terms of explaining image classifiers, with the aim of including negative results and insights we have gained. The pape…

cs.AI2016

A Meta-Analysis of the Anomaly Detection Problem

Andrew Emmott, Shubhomoy Das, Thomas Dietterich +2

This article provides a thorough meta-analysis of the anomaly detection problem. To accomplish this we first identify approaches to benchmarking anomaly detection algorithms across…

cs.RO2026

Learning Multi-Modal Whole-Body Control for Real-World Humanoid Robots

Pranay Dugar, Aayam Shrestha, Fangzhou Yu +2

A major challenge in humanoid robotics is designing a unified interface for commanding diverse whole-body behaviors, from precise footstep sequences to partial-body mimicry and joy…

cs.LG2025

DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs

Aayam Shrestha, Stefan Lee, Prasad Tadepalli +1

We study an approach to offline reinforcement learning (RL) based on optimally solving finitely-represented MDPs derived from a static dataset of experience. This approach can be a…

cs.RO2021

Learning Spring Mass Locomotion: Guiding Policies with a Reduced-Order Model

Kevin Green, Yesh Godse, Jeremy Dao +3

In this paper, we describe an approach to achieve dynamic legged locomotion on physical robots which combines existing methods for control with reinforcement learning. Specifically…

cs.LG2017

Incorporating Feedback into Tree-based Anomaly Detection

Shubhomoy Das, Weng-Keen Wong, Alan Fern +2

Anomaly detectors are often used to produce a ranked list of statistical anomalies, which are examined by human analysts in order to extract the actual anomalies of interest. Unfor…

cs.AI2025

WOFOSTGym: A Crop Simulator for Learning Annual and Perennial Crop Management Strategies

William Solow, Sandhya Saisubramanian, Alan Fern

We introduce WOFOSTGym, a novel crop simulation environment designed to train reinforcement learning (RL) agents to optimize agromanagement decisions for annual and perennial crops…

cs.DB2014

Representation Independent Analytics Over Structured Data

Yodsawalai Chodpathumwan, Jose Picado, Arash Termehchy +2

Database analytics algorithms leverage quantifiable structural properties of the data to predict interesting concepts and relationships. The same information, however, can be repre…

cs.LG2023

Multi-Task Learning for Budbreak Prediction

Aseem Saxena, Paola Pesantez-Cabrera, Rohan Ballapragada +2

Grapevine budbreak is a key phenological stage of seasonal development, which serves as a signal for the onset of active growth. This is also when grape plants are most vulnerable…

cs.LG2012

Output Space Search for Structured Prediction

Janardhan Rao Doppa, Alan Fern, Prasad Tadepalli

We consider a framework for structured prediction based on search in the space of complete structured outputs. Given a structured input, an output is produced by running a time-bou…

cs.RO2026

Simulator Adaptation for Sim-to-Real Learning of Legged Locomotion via Proprioceptive Distribution Matching

Jeremy Dao, Alan Fern

Simulation trained legged locomotion policies often exhibit performance loss on hardware due to dynamics discrepancies between the simulator and the real world, highlighting the ne…

cs.RO2023

Sim-to-Real Learning for Humanoid Box Loco-Manipulation

Jeremy Dao, Helei Duan, Alan Fern

In this work we propose a learning-based approach to box loco-manipulation for a humanoid robot. This is a particularly challenging problem due to the need for whole-body coordinat…

cs.PL2011

Adaptation-Based Programming in Haskell

Tim Bauer, Martin Erwig, Alan Fern +1

We present an embedded DSL to support adaptation-based programming (ABP) in Haskell. ABP is an abstract model for defining adaptive values, called adaptives, which adapt in respons…

cs.DB2017

Schema Independent Relational Learning

Jose Picado, Arash Termehchy, Alan Fern +1

Learning novel concepts and relations from relational databases is an important problem with many applications in database systems and machine learning. Relational learning algorit…

cs.AI2012

Inductive Policy Selection for First-Order MDPs

Sung Wook Yoon, Alan Fern, Robert Givan

We select policies for large Markov Decision Processes (MDPs) with compact first-order representations. We find policies that generalize well as the number of objects in the domain…

cs.AI2026

Grounding Vision and Language to 3D Masks for Long-Horizon Box Rearrangement

Ashish Malik, Caleb Lowe, Aayam Shrestha +3

We study long-horizon planning in 3D environments from under-specified natural-language goals using only visual observations, focusing on multi-step 3D box rearrangement tasks. Exi…

cs.LG2018

Interactive Naming for Explaining Deep Neural Networks: A Formative Study

Mandana Hamidi-Haines, Zhongang Qi, Alan Fern +2

We consider the problem of explaining the decisions of deep neural networks for image recognition in terms of human-recognizable visual concepts. In particular, given a test set of…

cs.RO2022

Dynamic Bipedal Maneuvers through Sim-to-Real Reinforcement Learning

Fangzhou Yu, Ryan Batke, Jeremy Dao +3

For legged robots to match the athletic capabilities of humans and animals, they must not only produce robust periodic walking and running, but also seamlessly switch between nomin…

cs.RO2025

Learning Decentralized Multi-Biped Control for Payload Transport

Bikram Pandit, Ashutosh Gupta, Mohitvishnu S. Gadde +5

Payload transport over flat terrain via multi-wheel robot carriers is well-understood, highly effective, and configurable. In this paper, our goal is to provide similar effectivene…