11 citations · 25 across the 14 of their papers we have counts for
11 papers · 1 filter
Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning
Meraj Mammadov, Pedro Zuidberg Dos Martires, Johannes Andreas Stork
Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex and high-dimensional observati…
APC-RL: Exceeding Data-Driven Behavior Priors with Adaptive Policy Composition
Finn Rietz, Pedro Zuidberg dos Martires, Johannes Andreas Stork
Incorporating demonstration data into reinforcement learning (RL) can greatly accelerate learning, but existing approaches often assume demonstrations are optimal and fully aligned…
KEA: Keeping Exploration Alive by Proactively Coordinating Exploration Strategies
Shih-Min Yang, Martin Magnusson, Johannes A. Stork +1
Soft Actor-Critic (SAC) has achieved notable success in continuous control tasks but struggles in sparse reward settings, where infrequent rewards make efficient exploration challe…
On the Effects of Irrelevant Variables in Treatment Effect Estimation with Deep Disentanglement
Ahmad Saeed Khan, Erik Schaffernicht, Johannes Andreas Stork
Estimating treatment effects from observational data is paramount in healthcare, education, and economics, but current deep disentanglement-based methods to address selection bias…
Learning Solutions of Stochastic Optimization Problems with Bayesian Neural Networks
Alan A. Lahoud, Erik Schaffernicht, Johannes A. Stork
Mathematical solvers use parametrized Optimization Problems (OPs) as inputs to yield optimal decisions. In many real-world settings, some of these parameters are unknown or uncerta…
DataSP: A Differential All-to-All Shortest Path Algorithm for Learning Costs and Predicting Paths with Context
Alan A. Lahoud, Erik Schaffernicht, Johannes A. Stork
Learning latent costs of transitions on graphs from trajectories demonstrations under various contextual features is challenging but useful for path planning. Yet, existing methods…