Inverse Contextual Bandits: Learning How Behavior Evolves over Time
arXiv:2107.06317
Abstract
Understanding a decision-maker's priorities by observing their behavior is critical for transparency and accountability in decision processes, such as in healthcare. Though conventional approaches to policy learning almost invariably assume stationarity in behavior, this is hardly true in practice: Medical practice is constantly evolving as clinical professionals fine-tune their knowledge over time. For instance, as the medical community's understanding of organ transplantations has progressed over the years, a pertinent question is: How have actual organ allocation policies been evolving? To give an answer, we desire a policy learning method that provides interpretable representations of decision-making, in particular capturing an agent's non-stationary knowledge of the world, as well as operating in an offline manner. First, we model the evolving behavior of decision-makers in terms of contextual bandits, and formalize the problem of Inverse Contextual Bandits (ICB). Second, we propose two concrete algorithms as solutions, learning parametric and nonparametric representations of an agent's behavior. Finally, using both real and simulated data for liver transplantations, we illustrate the applicability and explainability of our method, as well as benchmarking and validating its accuracy.
In Proceedings of the 39th International Conference on Machine Learning
References in corpus (12)
- Model-based Adversarial Imitation Learning
- What Did You Think Would Happen? Explaining Agent Behaviour Through Intended Outcomes
- Deep Bayesian Reward Learning from Preferences
- -GAIL: Learning -Divergence for Generative Adversarial Imitation Learning
- Learning a Multi-Modal Policy via Imitating Demonstrations with Mixed Behaviors
- Inverse Active Sensing: Modeling and Understanding Timely Decision-Making
- Explaining by Imitating: Understanding Decisions by Interpretable Policy Learning
- Inverse Decision Modeling: Learning Interpretable Representations of Behavior
- Efficient Exploration of Reward Functions in Inverse Reinforcement Learning via Bayesian Optimization
- Inverse Reinforcement Learning from a Gradient-based Learner
- Learning from an Exploring Demonstrator: Optimal Reward Estimation for Bandits
- Am I Building a White Box Agent or Interpreting a Black Box Agent?