Visualizing and Understanding Atari Agents
arXiv:1711.00138
Abstract
While deep reinforcement learning (deep RL) agents are effective at maximizing rewards, it is often unclear what strategies they use to do so. In this paper, we take a step toward explaining deep RL agents through a case study using Atari 2600 environments. In particular, we focus on using saliency maps to understand how an agent learns and executes a policy. We introduce a method for generating useful saliency maps and use it to show 1) what strong agents attend to, 2) whether agents are making decisions for the right or wrong reasons, and 3) how agents evolve during learning. We also test our method on non-expert human subjects and find that it improves their ability to reason about these agents. Overall, our results show that saliency information can provide significant insight into an RL agent's decisions and learning behavior.
ICML 2018 conference paper. Code: https://github.com/greydanus/visualize_atari Blog: https://greydanus.github.io/2017/11/01/visualize-atari/
Cited by in corpus (51)
- Scalable agent alignment via reward modeling: a research direction
- Interestingness Elements for Explainable Reinforcement Learning: Understanding Agents' Capabilities and Limitations
- Acquisition of Chess Knowledge in AlphaZero
- ChainerRL: A Deep Reinforcement Learning Library
- Counterfactual State Explanations for Reinforcement Learning Agents via Generative Deep Learning
- Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency Maps
- Restricting the Flow: Information Bottlenecks for Attribution
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
- Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations
- Towards falsifiable interpretability research
- Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
- Automatic Discovery of Interpretable Planning Strategies
- Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based Games
- The Emerging Landscape of Explainable AI Planning and Decision Making
- Observational Overfitting in Reinforcement Learning
- Counterfactual States for Atari Agents via Generative Deep Learning
- SECANT: Self-Expert Cloning for Zero-Shot Generalization of Visual Policies
- Learn to Interpret Atari Agents
- An initial attempt of combining visual selective attention with deep reinforcement learning
- Observation Space Matters: Benchmark and Optimization Algorithm
- Contrastive Explanations for Reinforcement Learning via Embedded Self Predictions
- Snooping Attacks on Deep Reinforcement Learning
- LIMEcraft: Handcrafted superpixel selection and inspection for Visual eXplanations
- Learning Symbolic Rules for Interpretable Deep Reinforcement Learning
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations
- A Game-Theoretic Taxonomy of Visual Concepts in DNNs
- Counterfactual Explanations in Sequential Decision Making Under Uncertainty
- The Automated Inspection of Opaque Liquid Vaccines
- Interactive Visualization for Debugging RL
- Domain Adaptation In Reinforcement Learning Via Latent Unified State Representation
- Flow-based Intrinsic Curiosity Module
- Understanding Learned Reward Functions
- Towards Robust Explanations for Deep Neural Networks
- Identifying Reasoning Flaws in Planning-Based RL Using Tree Explanations
- Are Gradient-based Saliency Maps Useful in Deep Reinforcement Learning?
- Distributed Reinforcement Learning of Targeted Grasping with Active Vision for Mobile Manipulators
- Mixture of Step Returns in Bootstrapped DQN
- Visual Diagnostics for Deep Reinforcement Learning Policy Development
- The Principle of Unchanged Optimality in Reinforcement Learning Generalization
- RoCUS: Robot Controller Understanding via Sampling
- Ranking Policy Decisions
- Why? Why not? When? Visual Explanations of Agent Behavior in Reinforcement Learning
- Generalized Constraints as A New Mathematical Problem in Artificial Intelligence: A Review and Perspective
- Visual Explanation using Attention Mechanism in Actor-Critic-based Deep Reinforcement Learning
- Vizarel: A System to Help Better Understand RL Agents
- Explanation of Reinforcement Learning Model in Dynamic Multi-Agent System
- RL agents Implicitly Learning Human Preferences
- Reconstructing Actions To Explain Deep Reinforcement Learning
- Sidekick Policy Learning for Active Visual Exploration