most citedAdapting Quality Assurance to Adaptive Systems: The Scenario Coevolution Paradigm

8 citations · 12 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG20204 cited

SAT-MARL: Specification Aware Training in Multi-Agent Reinforcement Learning

Fabian Ritz, Thomy Phan, Robert Müller +8

A characteristic of reinforcement learning is the ability to develop unforeseen strategies when solving problems. While such strategies sometimes yield superior performance, they m…

cs.LG2020

Policy Entropy for Out-of-Distribution Classification

Andreas Sedlmeier, Robert Müller, Steffen Illium +1

One critical prerequisite for the deployment of reinforcement learning systems in the real world is the ability to reliably detect situations on which the agent was not trained. Su…

cs.LG2020

Trajectory annotation using sequences of spatial perception

Sebastian Feld, Steffen Illium, Andreas Sedlmeier +1

In the near future, more and more machines will perform tasks in the vicinity of human spaces or support them directly in their spatially bound activities. In order to simplify the…

cs.LG2020

Bayesian Surprise in Indoor Environments

Sebastian Feld, Andreas Sedlmeier, Markus Friedrich +2

This paper proposes a novel method to identify unexpected structures in 2D floor plans using the concept of Bayesian Surprise. Taking into account that a person's expectation is an…

cs.LG2019

Uncertainty-Based Out-of-Distribution Classification in Deep Reinforcement Learning

Andreas Sedlmeier, Thomas Gabor, Thomy Phan +2

Robustness to out-of-distribution (OOD) data is an important goal in building reliable machine learning systems. Especially in autonomous systems, wrong predictions for OOD inputs…

cs.LG2019

Uncertainty-Based Out-of-Distribution Detection in Deep Reinforcement Learning

Andreas Sedlmeier, Thomas Gabor, Thomy Phan +2

We consider the problem of detecting out-of-distribution (OOD) samples in deep reinforcement learning. In a value based reinforcement learning setting, we propose to use uncertaint…