Playing Atari with Deep Reinforcement Learning
arXiv:1312.5602
Abstract
We present the first deep learning model to successfully learn control policies directly from high-dimensional sensory input using reinforcement learning. The model is a convolutional neural network, trained with a variant of Q-learning, whose input is raw pixels and whose output is a value function estimating future rewards. We apply our method to seven Atari 2600 games from the Arcade Learning Environment, with no adjustment of the architecture or learning algorithm. We find that it outperforms all previous approaches on six of the games and surpasses a human expert on three of them.
NIPS Deep Learning Workshop 2013
Cited by in corpus (1518)
- Deep Learning in Neural Networks: An Overview
- Continuous control with deep reinforcement learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Towards A Rigorous Science of Interpretable Machine Learning
- Trust Region Policy Optimization
- Ensemble deep learning: A review
- Prioritized Experience Replay
- Soft Actor-Critic Algorithms and Applications
- Asynchronous Methods for Deep Reinforcement Learning
- End-to-End Training of Deep Visuomotor Policies
- Learning Combinatorial Optimization Algorithms over Graphs
- Dota 2 with Large Scale Deep Reinforcement Learning
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Deep Reinforcement Learning for Cyber Security
- Regularized Deep Networks in Intelligent Transportation Systems: A Taxonomy and a Case Study
- Artificial Neural Networks trained through Deep Reinforcement Learning discover control strategies for active flow control
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Unity: A General Platform for Intelligent Agents
- Conservative Q-Learning for Offline Reinforcement Learning
- Deep Neural Networks for Bot Detection
- Deep reinforcement learning from human preferences
- A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers
- Decision Transformer: Reinforcement Learning via Sequence Modeling
- Hamiltonian Neural Networks
- Semi-supervised Deep Reinforcement Learning in Support of IoT and Smart City Services
- Action-Conditional Video Prediction using Deep Networks in Atari Games
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- Deep Reinforcement Learning for Page-wise Recommendations
- Reinforcement Learning with Deep Energy-Based Policies
- Deep Reinforcement Learning for Dialogue Generation
- Recommendations with Negative Feedback via Pairwise Deep Reinforcement Learning
- Massively Parallel Methods for Deep Reinforcement Learning
- Model-Based Reinforcement Learning for Atari
- Advances and Challenges in Conversational Recommender Systems: A Survey
- Learning to schedule job-shop problems: Representation and policy learning using graph neural network and reinforcement learning
- On Tiny Episodic Memories in Continual Learning
- Machine Learning for Wireless Communications in the Internet of Things: A Comprehensive Survey
- Split Computing and Early Exiting for Deep Learning Applications: Survey and Research Challenges
- Deep Reinforcement Learning for 5G Networks: Joint Beamforming, Power Control, and Interference Coordination
- Exponential expressivity in deep neural networks through transient chaos
- Maximum Entropy Deep Inverse Reinforcement Learning
- Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis
- Evaluation of deep learning models for multi-step ahead time series prediction
- Deep Reinforcement Learning meets Graph Neural Networks: exploring a routing optimization use case
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Model compression via distillation and quantization
- Behavior Regularized Offline Reinforcement Learning
- Exploration by Random Network Distillation
- Benchmarking Model-Based Reinforcement Learning
- A Study on Overfitting in Deep Reinforcement Learning
- Uncertainty-Aware Reinforcement Learning for Collision Avoidance
- Explanation in Human-AI Systems: A Literature Meta-Review, Synopsis of Key Ideas and Publications, and Bibliography for Explainable AI
- An Application of Deep Reinforcement Learning to Algorithmic Trading
- Distributive Dynamic Spectrum Access through Deep Reinforcement Learning: A Reservoir Computing Based Approach
- Deep Models Under the GAN: Information Leakage from Collaborative Deep Learning
- Robust active flow control over a range of Reynolds numbers using an artificial neural network trained through deep reinforcement learning
- A Survey of Machine Learning for Computer Architecture and Systems
- Physics Informed Neural Networks for Control Oriented Thermal Modeling of Buildings
- An Efficient Graph Convolutional Network Technique for the Travelling Salesman Problem
- Intelligent Power Control for Spectrum Sharing in Cognitive Radios: A Deep Reinforcement Learning Approach
- A Deep Reinforcement Learning Chatbot
- Differentiable MPC for End-to-end Planning and Control
- Quantifying Generalization in Reinforcement Learning
- Learning to Evade Static PE Machine Learning Malware Models via Reinforcement Learning
- Fathom: Reference Workloads for Modern Deep Learning Methods
- Ablation Studies in Artificial Neural Networks
- Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
- Deep Q-Learning for Same-Day Delivery with Vehicles and Drones
- MLPerf Training Benchmark
- Learning Combinatorial Optimization on Graphs: A Survey with Applications to Networking
- Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning
- Causal Confusion in Imitation Learning
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Learning to Perform Local Rewriting for Combinatorial Optimization
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Transfer Learning in Deep Reinforcement Learning: A Survey
- Deep Learning for Portfolio Optimization
- Multi-agent Reinforcement Learning for Cooperative Lane Changing of Connected and Autonomous Vehicles in Mixed Traffic
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Deep Reinforcement Learning for Process Control: A Primer for Beginners
- Physics-informed neural networks for the shallow-water equations on the sphere
- Deep Deterministic Policy Gradient for Urban Traffic Light Control
- End-to-End Deep Reinforcement Learning for Lane Keeping Assist
- Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
- DeepArchitect: Automatically Designing and Training Deep Architectures
- Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- The Modern Mathematics of Deep Learning
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
- Relational Neural Expectation Maximization: Unsupervised Discovery of Objects and their Interactions
- Battery-constrained Federated Edge Learning in UAV-enabled IoT for B5G/6G Networks
- Neural Interactive Collaborative Filtering
- Generative Design by Reinforcement Learning: Enhancing the Diversity of Topology Optimization Designs
- Tianshou: a Highly Modularized Deep Reinforcement Learning Library
- Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents
- Quantum Neuron: an elementary building block for machine learning on quantum computers
- DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections
- DeepPath: A Reinforcement Learning Method for Knowledge Graph Reasoning
- A review on deep reinforcement learning for fluid mechanics: an update
- Reinforcement Learning Neural Turing Machines - Revised
- A review on Deep Reinforcement Learning for Fluid Mechanics
- Deep Reinforcement Learning for Combinatorial Optimization: Covering Salesman Problems
- GraphVite: A High-Performance CPU-GPU Hybrid System for Node Embedding
- Deep Reinforcement Learning for List-wise Recommendations
- What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
- Review: Deep Learning in Electron Microscopy
- Learning Visual Predictive Models of Physics for Playing Billiards
- RUDDER: Return Decomposition for Delayed Rewards
- Deep Reinforcement Learning and the Deadly Triad
- SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving
- Model-Based Meta-Reinforcement Learning for Flight with Suspended Payloads
- Deep Reinforcement Learning for Visual Object Tracking in Videos
- Implicit Generation and Generalization in Energy-Based Models
- An Optimistic Perspective on Offline Reinforcement Learning
- Practical Black-Box Attacks against Machine Learning
- Deep Reinforcement Learning for Black-Box Testing of Android Apps
- Deep Reinforcement Learning Optimizes Graphene Nanopores for Efficient Desalination
- Deep Reinforcement Learning for Real-Time Optimization of Pumps in Water Distribution Systems
- Loss is its own Reward: Self-Supervision for Reinforcement Learning
- Measuring the Algorithmic Efficiency of Neural Networks
- Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning
- CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms
- Meta-Learning Update Rules for Unsupervised Representation Learning
- Learning to Optimize Join Queries With Deep Reinforcement Learning
- A systematic review of fuzzing based on machine learning techniques
- Tree-Structured Reinforcement Learning for Sequential Object Localization
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
- Self-Driving Car Steering Angle Prediction Based on Image Recognition
- Knowledge Transfer between Buildings for Seismic Damage Diagnosis through Adversarial Learning
- Projection-Based Constrained Policy Optimization
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Fusion of Model-free Reinforcement Learning with Microgrid Control: Review and Vision
- Towards Cognitive Exploration through Deep Reinforcement Learning for Mobile Robots
- Quantum circuit optimization with deep reinforcement learning
- DeepRobust: A PyTorch Library for Adversarial Attacks and Defenses
- On Improving Deep Reinforcement Learning for POMDPs
- A Lyapunov-based Approach to Safe Reinforcement Learning
- Jointly Learning to Recommend and Advertise
- Exploring Model-based Planning with Policy Networks
- Playing Atari Games with Deep Reinforcement Learning and Human Checkpoint Replay
- Gradient based sample selection for online continual learning
- Machine Learning in a data-limited regime: Augmenting experiments with synthetic data uncovers order in crumpled sheets
- DRLinFluids -- An open-source python platform of coupling Deep Reinforcement Learning and OpenFOAM
- CHALET: Cornell House Agent Learning Environment
- ChainerRL: A Deep Reinforcement Learning Library
- Unrestricted Adversarial Examples
- Dynamics-Aware Unsupervised Discovery of Skills
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- Using Recurrent Neural Networks to Optimize Dynamical Decoupling for Quantum Memory
- Intelligent Residential Energy Management System using Deep Reinforcement Learning
- Model-Based Reinforcement Learning with Value-Targeted Regression
- Learning Heuristics over Large Graphs via Deep Reinforcement Learning
- Latent Space Policies for Hierarchical Reinforcement Learning
- Giraffe: Using Deep Reinforcement Learning to Play Chess
- 3D Simulation for Robot Arm Control with Deep Q-Learning
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
- Active One-shot Learning
- Open-Sourced Reinforcement Learning Environments for Surgical Robotics
- Deep Reinforcement learning for real autonomous mobile robot navigation in indoor environments
- Combining Deep Reinforcement Learning and Safety Based Control for Autonomous Driving
- A Reinforcement Learning-based Economic Model Predictive Control Framework for Autonomous Operation of Chemical Reactors
- Neural Network Based Reinforcement Learning for Audio-Visual Gaze Control in Human-Robot Interaction
- Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- DRLE: Decentralized Reinforcement Learning at the Edge for Traffic Light Control in the IoV
- Self-Supervised Policy Adaptation during Deployment
- Autonomous Penetration Testing using Reinforcement Learning
- Actor-Critic Method for High Dimensional Static Hamilton--Jacobi--Bellman Partial Differential Equations based on Neural Networks
- Solving the Order Batching and Sequencing Problem using Deep Reinforcement Learning
- A Derivative-Free Method for Solving Elliptic Partial Differential Equations with Deep Neural Networks
- Sobolev Training for Neural Networks
- Adaptive Behavior Generation for Autonomous Driving using Deep Reinforcement Learning with Compact Semantic States
- Deep Hierarchical Reinforcement Learning Algorithm in Partially Observable Markov Decision Processes
- A Genetic Programming Approach to Designing Convolutional Neural Network Architectures
- Distributed Deep Q-Learning
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- Multi-Task Reinforcement Learning with Soft Modularization
- On Characterizing the Capacity of Neural Networks using Algebraic Topology
- Dynamic Weights in Multi-Objective Deep Reinforcement Learning
- Transfer Learning for Related Reinforcement Learning Tasks via Image-to-Image Translation
- Strategic Dialogue Management via Deep Reinforcement Learning
- A Self-adaptive SAC-PID Control Approach based on Reinforcement Learning for Mobile Robots
- Effective control of two-dimensional Rayleigh--Bénard convection: invariant multi-agent reinforcement learning is all you need
- The Uncertainty Bellman Equation and Exploration
- Dropout as a Bayesian Approximation: Appendix
- Graying the black box: Understanding DQNs
- Multiagent Soft Q-Learning
- Unsupervised Video Object Segmentation for Deep Reinforcement Learning
- Learning What Data to Learn
- Reinforcement learning for optimization of variational quantum circuit architectures
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- Exposure: A White-Box Photo Post-Processing Framework
- First Order Constrained Optimization in Policy Space
- Automatic Data Augmentation for Generalization in Deep Reinforcement Learning
- Characterizing Attacks on Deep Reinforcement Learning
- Training Neural Networks Using Features Replay
- rlpyt: A Research Code Base for Deep Reinforcement Learning in PyTorch
- The Transformer Network for the Traveling Salesman Problem
- Learning End-to-end Multimodal Sensor Policies for Autonomous Navigation
- A reinforcement learning approach to rare trajectory sampling
- GenDICE: Generalized Offline Estimation of Stationary Values
- Dealing with Sparse Rewards in Reinforcement Learning
- Efficient Processing of Deep Neural Networks: A Tutorial and Survey
- DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback
- Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure
- Mid-Level Visual Representations Improve Generalization and Sample Efficiency for Learning Visuomotor Policies
- Global Optimization of Gaussian processes
- Faster Fuzzing: Reinitialization with Deep Neural Models
- A Reinforcement Learning Environment For Job-Shop Scheduling
- Reinforcement Learning with Combinatorial Actions: An Application to Vehicle Routing
- Superposition of many models into one
- Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening
- A Review of Designs and Applications of Echo State Networks
- Playing Doom with SLAM-Augmented Deep Reinforcement Learning
- Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
- A Survey of Optimization Methods from a Machine Learning Perspective
- Deep Reinforcement Learning for Autonomous Driving
- Effective Diversity in Population Based Reinforcement Learning
- Reinforcement Learning Based Vehicle-cell Association Algorithm for Highly Mobile Millimeter Wave Communication
- Robust Distant Supervision Relation Extraction via Deep Reinforcement Learning
- Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation
- Neural Certificates for Safe Control Policies
- MazeBase: A Sandbox for Learning from Games
- Gym-ANM: Reinforcement Learning Environments for Active Network Management Tasks in Electricity Distribution Systems
- The Case for Automatic Database Administration using Deep Reinforcement Learning
- Learning Bilingual Word Representations by Marginalizing Alignments
- OnSlicing: Online End-to-End Network Slicing with Reinforcement Learning
- Embodied Visual Navigation with Automatic Curriculum Learning in Real Environments
- PoPS: Policy Pruning and Shrinking for Deep Reinforcement Learning
- Reinforcement Learning with Prototypical Representations
- CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
- Deep Reinforcement Learning of Cell Movement in the Early Stage of C. elegans Embryogenesis
- Domain Generalization with MixStyle
- Deep Q-Network Based Decision Making for Autonomous Driving
- Joint Modeling of Dense and Incomplete Trajectories for Citywide Traffic Volume Inference
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound
- Improving Electron Micrograph Signal-to-Noise with an Atrous Convolutional Encoder-Decoder
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- Adversarial Attacks and Defenses on Graphs: A Review, A Tool and Empirical Studies
- A Framework for Automated Cellular Network Tuning with Reinforcement Learning
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
- A DRL-based Multiagent Cooperative Control Framework for CAV Networks: a Graphic Convolution Q Network
- Have You Stolen My Model? Evasion Attacks Against Deep Neural Network Watermarking Techniques
- Reward learning from human preferences and demonstrations in Atari
- DRLViz: Understanding Decisions and Memory in Deep Reinforcement Learning
- Deep Learning Based Chatbot Models
- Constrained Attractor Selection Using Deep Reinforcement Learning
- Deep Implicit Coordination Graphs for Multi-agent Reinforcement Learning
- Towards a Common Implementation of Reinforcement Learning for Multiple Robotic Tasks
- DeepTraffic: Crowdsourced Hyperparameter Tuning of Deep Reinforcement Learning Systems for Multi-Agent Dense Traffic Navigation
- ReLMoGen: Leveraging Motion Generation in Reinforcement Learning for Mobile Manipulation
- Transforming Cooling Optimization for Green Data Center via Deep Reinforcement Learning
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- LIFT: Reinforcement Learning in Computer Systems by Learning From Demonstrations
- Multi-Hop Knowledge Graph Reasoning with Reward Shaping
- A Neural Transducer
- Improving Efficiency of Training a Virtual Treatment Planner Network via Knowledge-guided Deep Reinforcement Learning for Intelligent Automatic Treatment Planning of Radiotherapy
- Memory Augmented Control Networks
- Deep Q-Learning for Self-Organizing Networks Fault Management and Radio Performance Improvement
- Deep Multi-agent Reinforcement Learning for Highway On-Ramp Merging in Mixed Traffic
- Rethinking the Implementation Tricks and Monotonicity Constraint in Cooperative Multi-Agent Reinforcement Learning
- Wasserstein Robust Reinforcement Learning
- MLGO: a Machine Learning Guided Compiler Optimizations Framework
- Reinforcement learning for optimal error correction of toric codes
- Deconfounding Reinforcement Learning in Observational Settings
- Deep Multi-Agent Reinforcement Learning with Relevance Graphs
- A Benchmarking Environment for Reinforcement Learning Based Task Oriented Dialogue Management
- Predicting Game Difficulty and Churn Without Players
- Prior-Knowledge and Attention-based Meta-Learning for Few-Shot Learning
- Meta-learning of Sequential Strategies
- State of the Art Control of Atari Games Using Shallow Reinforcement Learning
- Automated Database Indexing using Model-free Reinforcement Learning
- On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models
- Prioritized Sequence Experience Replay
- Neural Graph Evolution: Towards Efficient Automatic Robot Design
- Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain Challenge
- Deterministic Implementations for Reproducibility in Deep Reinforcement Learning
- Hierarchical Decision Making by Generating and Following Natural Language Instructions
- Beating Atari with Natural Language Guided Reinforcement Learning
- Understanding Domain Randomization for Sim-to-real Transfer
- Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP
- ConvLab: Multi-Domain End-to-End Dialog System Platform
- A Review of Machine Learning Applications in Fuzzing
- A Non-Technical Survey on Deep Convolutional Neural Network Architectures
- CausalWorld: A Robotic Manipulation Benchmark for Causal Structure and Transfer Learning
- Imitation Learning for Non-Autoregressive Neural Machine Translation
- Natural Environment Benchmarks for Reinforcement Learning
- Mapping State Space using Landmarks for Universal Goal Reaching
- Model-free Deep Reinforcement Learning for Urban Autonomous Driving
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation
- Controlled Online Optimization Learning (COOL): Finding the ground state of spin Hamiltonians with reinforcement learning
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder Dimension
- A non-cooperative meta-modeling game for automated third-party calibrating, validating, and falsifying constitutive laws with parallelized adversarial attacks
- Parameterized Reinforcement Learning for Optical System Optimization
- Diff-DAC: Distributed Actor-Critic for Average Multitask Deep Reinforcement Learning
- Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient Algorithms
- Laplacian Smoothing Gradient Descent
- SCC: an efficient deep reinforcement learning agent mastering the game of StarCraft II
- Deep Reinforcement Learning based Resource Allocation for V2V Communications
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- Combating the Compounding-Error Problem with a Multi-step Model
- Continual Learning with Node-Importance based Adaptive Group Sparse Regularization
- Towards Guaranteed Safety Assurance of Automated Driving Systems with Scenario Sampling: An Invariant Set Perspective (Extended Version)
- Optimal Control Via Neural Networks: A Convex Approach
- Deep reinforcement learning for universal quantum state preparation via dynamic pulse control
- Succinct and Robust Multi-Agent Communication With Temporal Message Control
- Autonomous Air Traffic Controller: A Deep Multi-Agent Reinforcement Learning Approach
- Explore, Exploit or Listen: Combining Human Feedback and Policy Model to Speed up Deep Reinforcement Learning in 3D Worlds
- Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
- Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
- Exponentially Increasing the Capacity-to-Computation Ratio for Conditional Computation in Deep Learning
- Unsupervised State Representation Learning in Atari
- A Deep Multi-Agent Reinforcement Learning Approach to Autonomous Separation Assurance
- Constructing Parsimonious Analytic Models for Dynamic Systems via Symbolic Regression
- FuzzerGym: A Competitive Framework for Fuzzing and Learning
- A Survey of Embodied AI: From Simulators to Research Tasks
- NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning
- MALib: A Parallel Framework for Population-based Multi-agent Reinforcement Learning
- Autonomous and cooperative design of the monitor positions for a team of UAVs to maximize the quantity and quality of detected objects
- Online Deep Reinforcement Learning for Autonomous UAV Navigation and Exploration of Outdoor Environments
- Approximate Inference with Amortised MCMC
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- Network Environment Design for Autonomous Cyberdefense
- Learning Robotic Manipulation of Granular Media
- Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control
- Reinforced active learning for image segmentation
- TriFinger: An Open-Source Robot for Learning Dexterity
- Learning to Explore with Meta-Policy Gradient
- Learning Functionally Decomposed Hierarchies for Continuous Control Tasks with Path Planning
- Interpretable End-to-end Urban Autonomous Driving with Latent Deep Reinforcement Learning
- Lifelong Object Detection
- Discriminative Particle Filter Reinforcement Learning for Complex Partial Observations
- Deep Reinforcement and InfoMax Learning
- Learning to Continuously Optimize Wireless Resource In Episodically Dynamic Environment
- Learning Simple Algorithms from Examples
- Benchmarks for Deep Off-Policy Evaluation
- DISCO: Influence Maximization Meets Network Embedding and Deep Learning
- Behavioral decision-making for urban autonomous driving in the presence of pedestrians using Deep Recurrent Q-Network
- ABIDES-Gym: Gym Environments for Multi-Agent Discrete Event Simulation and Application to Financial Markets
- Is Q-learning Provably Efficient?
- A Deep Q-Network for the Beer Game: A Deep Reinforcement Learning algorithm to Solve Inventory Optimization Problems
- Combining Neural Networks and Tree Search for Task and Motion Planning in Challenging Environments
- Federated Control with Hierarchical Multi-Agent Deep Reinforcement Learning
- Polygonal Building Segmentation by Frame Field Learning
- A Critical Investigation of Deep Reinforcement Learning for Navigation
- Fast Neural Network Verification via Shadow Prices
- Virtual-Taobao: Virtualizing Real-world Online Retail Environment for Reinforcement Learning
- Deep Reinforcement Learning for Unmanned Aerial Vehicle-Assisted Vehicular Networks
- Double Deep Q-Learning for Optimal Execution
- Replay in Deep Learning: Current Approaches and Missing Biological Elements
- Modular Deep Reinforcement Learning with Temporal Logic Specifications
- Probing the Theoretical and Computational Limits of Dissipative Design
- Correspondence between neuroevolution and gradient descent
- Enhancing the Insertion of NOP Instructions to Obfuscate Malware via Deep Reinforcement Learning
- Sample-Efficient Imitation Learning via Generative Adversarial Nets
- A Programmable Approach to Neural Network Compression
- Reinforcement Learning to Optimize Long-term User Engagement in Recommender Systems
- ScreenerNet: Learning Self-Paced Curriculum for Deep Neural Networks
- Deep Reinforcement Learning for Imbalanced Classification
- ACTRCE: Augmenting Experience via Teacher's Advice For Multi-Goal Reinforcement Learning
- CyGIL: A Cyber Gym for Training Autonomous Agents over Emulated Network Systems
- Graph Policy Gradients for Large Scale Robot Control
- Self-driving scale car trained by Deep reinforcement learning
- Stability-certified reinforcement learning: A control-theoretic perspective
- Cooperative Lane Changing via Deep Reinforcement Learning
- Reinforcement Learning on Variable Impedance Controller for High-Precision Robotic Assembly
- Interactively shaping robot behaviour with unlabeled human instructions
- Bilinear Classes: A Structural Framework for Provable Generalization in RL
- Dynamic Fusion Networks for Machine Reading Comprehension
- Toward automatic comparison of visualization techniques: Application to graph visualization
- Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
- Aggregating E-commerce Search Results from Heterogeneous Sources via Hierarchical Reinforcement Learning
- Learning to Group: A Bottom-Up Framework for 3D Part Discovery in Unseen Categories
- Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based Games
- Bipedal Walking Robot using Deep Deterministic Policy Gradient
- Millimeter Wave Communications with an Intelligent Reflector: Performance Optimization and Distributional Reinforcement Learning
- Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety
- Online Meta-Critic Learning for Off-Policy Actor-Critic Methods
- A Survey on Interactive Reinforcement Learning: Design Principles and Open Challenges
- Deep Reinforcement Learning for Dynamic Multichannel Access in Wireless Networks
- Discovering General-Purpose Active Learning Strategies
- Learning First-to-Spike Policies for Neuromorphic Control Using Policy Gradients
- Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes
- Multi-Agent Collaboration via Reward Attribution Decomposition
- Learning Deep Control Policies for Autonomous Aerial Vehicles with MPC-Guided Policy Search
- Learning Robust Dialog Policies in Noisy Environments
- PlanGAN: Model-based Planning With Sparse Rewards and Multiple Goals
- A Machine Learning Approach to Routing
- The problem with DDPG: understanding failures in deterministic environments with sparse rewards
- Distributionally Robust Reinforcement Learning
- Learning Robotic Manipulation through Visual Planning and Acting
- The Effectiveness of Memory Replay in Large Scale Continual Learning
- An Autonomous Free Airspace En-route Controller using Deep Reinforcement Learning Techniques
- Online Data Poisoning Attack
- Reconfigurable Intelligent Surface Enhanced Device-to-Device Communications
- CraftAssist: A Framework for Dialogue-enabled Interactive Agents
- Optimal Attacks on Reinforcement Learning Policies
- Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation
- Deep Robust Kalman Filter
- Exploring Unsupervised Pretraining and Sentence Structure Modelling for Winograd Schema Challenge
- Learning to Play No-Press Diplomacy with Best Response Policy Iteration
- Automated Lane Change Strategy using Proximal Policy Optimization-based Deep Reinforcement Learning
- Sparse Graphical Memory for Robust Planning
- A View on Deep Reinforcement Learning in System Optimization
- Building Safer Autonomous Agents by Leveraging Risky Driving Behavior Knowledge
- Strategies for Using Proximal Policy Optimization in Mobile Puzzle Games
- Defending Against Adversarial Examples with K-Nearest Neighbor
- Toward Simulating Environments in Reinforcement Learning Based Recommendations
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Deep reinforcement learning for time series: playing idealized trading games
- CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
- Applications of Deep Reinforcement Learning in Communications and Networking: A Survey
- Composing Meta-Policies for Autonomous Driving Using Hierarchical Deep Reinforcement Learning
- Reinforcement Learning for Education: Opportunities and Challenges
- Delay-Aware Multi-Agent Reinforcement Learning for Cooperative and Competitive Environments
- Observational Overfitting in Reinforcement Learning
- Using Reinforcement Learning with Partial Vehicle Detection for Intelligent Traffic Signal Control
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- SAMBA: Safe Model-Based & Active Reinforcement Learning
- DRLDO: A novel DRL based De-ObfuscationSystem for Defense against Metamorphic Malware
- Generalization in Reinforcement Learning by Soft Data Augmentation
- Deep Reinforcement Learning with Robust and Smooth Policy
- Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
- A Bi-Level Framework for Learning to Solve Combinatorial Optimization on Graphs
- Towards Cooperation in Sequential Prisoner's Dilemmas: a Deep Multiagent Reinforcement Learning Approach
- Sample-Efficient Reinforcement Learning via Counterfactual-Based Data Augmentation
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in Hanabi
- Partially Observable Markov Decision Process for Recommender Systems
- Reinforcement Learning with Quantum Variational Circuits
- Experience Replay Optimization
- Harnessing Structures for Value-Based Planning and Reinforcement Learning
- Learning to Factor Policies and Action-Value Functions: Factored Action Space Representations for Deep Reinforcement learning
- Deep Reinforcement Learning for Contact-Rich Skills Using Compliant Movement Primitives
- Scalable Centralized Deep Multi-Agent Reinforcement Learning via Policy Gradients
- Does AlphaGo actually play Go? Concerning the State Space of Artificial Intelligence
- AI-Based Autonomous Line Flow Control via Topology Adjustment for Maximizing Time-Series ATCs
- Contextual Imagined Goals for Self-Supervised Robotic Learning
- Reinforcement Learning with Perturbed Rewards
- Navigating Assistance System for Quadcopter with Deep Reinforcement Learning
- Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
- Recurrent Relational Networks
- Optimization Methods for Interpretable Differentiable Decision Trees in Reinforcement Learning
- Learning to Reason in Large Theories without Imitation
- Stellar Spectral Interpolation using Machine Learning
- Experience Replay Using Transition Sequences
- Flight Controller Synthesis Via Deep Reinforcement Learning
- RTFM: Generalising to Novel Environment Dynamics via Reading
- Dynamic Deep Neural Networks: Optimizing Accuracy-Efficiency Trade-offs by Selective Execution
- Modern Deep Reinforcement Learning Algorithms
- Deep Reinforcement Learning for Task Offloading in Mobile Edge Computing Systems
- Learning Nash Equilibrium for General-Sum Markov Games from Batch Data
- Deep Reinforcement Learning: Framework, Applications, and Embedded Implementations
- Model-based Adversarial Meta-Reinforcement Learning
- Robust Deep Reinforcement Learning through Adversarial Loss
- Decision-making at Unsignalized Intersection for Autonomous Vehicles: Left-turn Maneuver with Deep Reinforcement Learning
- Playing Atari with Hybrid Quantum-Classical Reinforcement Learning
- An application of data driven reward of deep reinforcement learning by dynamic mode decomposition in active flow control
- Hindsight Goal Ranking on Replay Buffer for Sparse Reward Environment
- MixStyle Neural Networks for Domain Generalization and Adaptation
- Quantum Continual Learning Overcoming Catastrophic Forgetting
- Estimation Error Correction in Deep Reinforcement Learning for Deterministic Actor-Critic Methods
- Adaptive Discretization in Online Reinforcement Learning
- Minimax Value Interval for Off-Policy Evaluation and Policy Optimization
- Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network
- ResearchDoom and CocoDoom: Learning Computer Vision with Games
- Constrained Model-based Reinforcement Learning with Robust Cross-Entropy Method
- Provably Efficient Online Hyperparameter Optimization with Population-Based Bandits
- A State Aggregation Approach for Solving Knapsack Problem with Deep Reinforcement Learning
- Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning
- Learning to Control in Metric Space with Optimal Regret
- SECANT: Self-Expert Cloning for Zero-Shot Generalization of Visual Policies
- Precision medicine as a control problem: Using simulation and deep reinforcement learning to discover adaptive, personalized multi-cytokine therapy for sepsis
- Multi-agent Hierarchical Reinforcement Learning with Dynamic Termination
- Obstacle Avoidance Using a Monocular Camera
- Stochastic Recursive Momentum for Policy Gradient Methods
- Towards Symbolic Reinforcement Learning with Common Sense
- A Short Survey On Memory Based Reinforcement Learning
- AirSim Drone Racing Lab
- Transferring Autonomous Driving Knowledge on Simulated and Real Intersections
- A Deep Reinforcement Learning Chatbot (Short Version)
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Tuning-free Plug-and-Play Proximal Algorithm for Inverse Imaging Problems
- Relating Graph Neural Networks to Structural Causal Models
- Efficient (Soft) Q-Learning for Text Generation with Limited Good Data
- Discretizing Continuous Action Space for On-Policy Optimization
- Continual Learning in Neural Networks
- Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
- Elements of Effective Deep Reinforcement Learning towards Tactical Driving Decision Making
- A Deep Generative Deconvolutional Image Model
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- Using Reinforcement Learning to find Efficient Qubit Routing Policies for Deployment in Near-term Quantum Computers
- Single Deep Counterfactual Regret Minimization
- Invariant Policy Optimization: Towards Stronger Generalization in Reinforcement Learning
- The Dynamics of Handwriting Improves the Automated Diagnosis of Dysgraphia
- CAQL: Continuous Action Q-Learning
- Learn to Interpret Atari Agents
- DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures
- Reinforcement Learning with A* and a Deep Heuristic
- Hedging using reinforcement learning: Contextual -Armed Bandit versus -learning
- Deep Reinforcement Learning for Chinese Zero pronoun Resolution
- Automated Driving Maneuvers under Interactive Environment based on Deep Reinforcement Learning
- Lane Change Decision-making through Deep Reinforcement Learning with Rule-based Constraints
- Gradient Band-based Adversarial Training for Generalized Attack Immunity of A3C Path Finding
- Deep Reinforcement Learning for Trading
- Online Robustness Training for Deep Reinforcement Learning
- Autonomous Intelligent Cyber-defense Agent (AICA) Reference Architecture. Release 2.0
- MAPPER: Multi-Agent Path Planning with Evolutionary Reinforcement Learning in Mixed Dynamic Environments
- Sample Efficient Reinforcement Learning via Low-Rank Matrix Estimation
- Fine-Grained Gap-Dependent Bounds for Tabular MDPs via Adaptive Multi-Step Bootstrap
- Reinforcement Learning-Driven Test Generation for Android GUI Applications using Formal Specifications
- Making Sense of Reinforcement Learning and Probabilistic Inference
- SMiRL: Surprise Minimizing Reinforcement Learning in Unstable Environments
- Learning Heuristic Search via Imitation
- PIC: Permutation Invariant Critic for Multi-Agent Deep Reinforcement Learning
- An Empirical Study of Representation Learning for Reinforcement Learning in Healthcare
- Model-based Reinforcement Learning for Service Mesh Fault Resiliency in a Web Application-level
- Solve Traveling Salesman Problem by Monte Carlo Tree Search and Deep Neural Network
- The Causal-Neural Connection: Expressiveness, Learnability, and Inference
- Do Offline Metrics Predict Online Performance in Recommender Systems?
- Top-K Off-Policy Correction for a REINFORCE Recommender System
- Provably Robust Blackbox Optimization for Reinforcement Learning
- HRL4IN: Hierarchical Reinforcement Learning for Interactive Navigation with Mobile Manipulators
- Ethical Artificial Intelligence
- Opportunistic Learning: Budgeted Cost-Sensitive Learning from Data Streams
- GreedyNAS: Towards Fast One-Shot NAS with Greedy Supernet
- Deep Reinforcement Learning for Clinical Decision Support: A Brief Survey
- Time-Varying Formation Controllers for Unmanned Aerial Vehicles Using Deep Reinforcement Learning
- Automated vehicle's behavior decision making using deep reinforcement learning and high-fidelity simulation environment
- Cooperative Perception with Deep Reinforcement Learning for Connected Vehicles
- Hyperbolic Deep Neural Networks: A Survey
- SpikePropamine: Differentiable Plasticity in Spiking Neural Networks
- Scalable Bayesian Inverse Reinforcement Learning
- Brain Inspired Cognitive Model with Attention for Self-Driving Cars
- Self-Imitation Learning via Generalized Lower Bound Q-learning
- Soft Hindsight Experience Replay
- A New Approach for Resource Scheduling with Deep Reinforcement Learning
- STMARL: A Spatio-Temporal Multi-Agent Reinforcement Learning Approach for Cooperative Traffic Light Control
- Analyzing Knowledge Transfer in Deep Q-Networks for Autonomously Handling Multiple Intersections
- Deep reinforcement learning approach to MIMO precoding problem: Optimality and Robustness
- Sample-Efficient Reinforcement Learning through Transfer and Architectural Priors
- Double Prioritized State Recycled Experience Replay
- GAN-powered Deep Distributional Reinforcement Learning for Resource Management in Network Slicing
- Offline Meta-Reinforcement Learning with Advantage Weighting
- Indoor Path Planning for an Unmanned Aerial Vehicle via Curriculum Learning
- Adversarial Imitation Learning via Random Search
- On the Utility of Model Learning in HRI
- Ablation of a Robot's Brain: Neural Networks Under a Knife
- Student-Initiated Action Advising via Advice Novelty
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning
- Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework
- Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning
- ORRB -- OpenAI Remote Rendering Backend
- Is Deep Reinforcement Learning Ready for Practical Applications in Healthcare? A Sensitivity Analysis of Duel-DDQN for Hemodynamic Management in Sepsis Patients
- Interference and Generalization in Temporal Difference Learning
- Deep Reinforcement Learning With Macro-Actions
- Map-based Multi-Policy Reinforcement Learning: Enhancing Adaptability of Robots by Deep Reinforcement Learning
- Diverse Auto-Curriculum is Critical for Successful Real-World Multiagent Learning Systems
- Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning
- SimpleDS: A Simple Deep Reinforcement Learning Dialogue System
- Why Build an Assistant in Minecraft?
- Self-supervised Learning of Distance Functions for Goal-Conditioned Reinforcement Learning
- Task-Agnostic Dynamics Priors for Deep Reinforcement Learning
- Snooping Attacks on Deep Reinforcement Learning
- A Neural Networks Committee for the Contextual Bandit Problem
- Divide-and-Conquer Reinforcement Learning
- Deep Q-Learning for Dynamic Reliability Aware NFV-Based Service Provisioning
- Run Away From your Teacher: Understanding BYOL by a Novel Self-Supervised Approach
- Faster and Safer Training by Embedding High-Level Knowledge into Deep Reinforcement Learning
- Cautious Reinforcement Learning via Distributional Risk in the Dual Domain
- Schedule Earth Observation satellites with Deep Reinforcement Learning
- Deep Reinforcement Learning in a Monetary Model
- Zeus: Efficiently Localizing Actions in Videos using Reinforcement Learning
- Transformer Based Reinforcement Learning For Games
- Emergence of Theory of Mind Collaboration in Multiagent Systems
- Robotic Surgery With Lean Reinforcement Learning
- Safe Deep Reinforcement Learning for Multi-Agent Systems with Continuous Action Spaces
- Improving Robustness of Reinforcement Learning for Power System Control with Adversarial Training
- Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian Sampling
- Microstructure Representation and Reconstruction of Heterogeneous Materials via Deep Belief Network for Computational Material Design
- Deep Active Learning for Dialogue Generation
- Reinforcement Learning-based Switching Controller for a Milliscale Robot in a Constrained Environment
- Symbolic Regression Methods for Reinforcement Learning
- Deep Reinforcement Learning for Sponsored Search Real-time Bidding
- Affordance-based Reinforcement Learning for Urban Driving
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Continuous Motion Planning with Temporal Logic Specifications using Deep Neural Networks
- Learning predictive representations in autonomous driving to improve deep reinforcement learning
- Optimizing Mixed Autonomy Traffic Flow With Decentralized Autonomous Vehicles and Multi-Agent RL
- Maximum Mutation Reinforcement Learning for Scalable Control
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- Adaptive Discretization for Model-Based Reinforcement Learning
- Offline Reinforcement Learning for Autonomous Driving with Safety and Exploration Enhancement
- Meta Reinforcement Learning with Task Embedding and Shared Policy
- Incorporating Relational Background Knowledge into Reinforcement Learning via Differentiable Inductive Logic Programming
- Deep Reinforcement Agent for Scheduling in HPC
- A Deep Reinforcement Learning Approach for Global Routing
- Gym-ANM: Open-source software to leverage reinforcement learning for power system management in research and education
- Deep Reinforcement Learning Based Volt-VAR Optimization in Smart Distribution Systems
- SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning
- Thinking While Moving: Deep Reinforcement Learning with Concurrent Control
- Semantic-Transferable Weakly-Supervised Endoscopic Lesions Segmentation
- Quadratic Q-network for Learning Continuous Control for Autonomous Vehicles
- Directed Exploration for Reinforcement Learning
- Smooth Exploration for Robotic Reinforcement Learning
- Learning to Plan Hierarchically from Curriculum
- Learning Adaptive Exploration Strategies in Dynamic Environments Through Informed Policy Regularization
- Recurrent Predictive State Policy Networks
- Deep Reinforcement Learning for Dynamic Urban Transportation Problems
- Automatic Testing With Reusable Adversarial Agents
- Toward an Automated Auction Framework for Wireless Federated Learning Services Market
- Reinforcement Learning in Rich-Observation MDPs using Spectral Methods
- Neural Stochastic Dual Dynamic Programming
- Safe Reinforcement Learning Using Robust Action Governor
- Noise-Robust End-to-End Quantum Control using Deep Autoregressive Policy Networks
- A Dual-Hormone Closed-Loop Delivery System for Type 1 Diabetes Using Deep Reinforcement Learning
- Offline Reinforcement Learning with Soft Behavior Regularization
- Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space
- Chance-Constrained Control with Lexicographic Deep Reinforcement Learning
- Assumed Density Filtering Q-learning
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL
- Deep Reinforcement Learning for Dynamic Treatment Regimes on Medical Registry Data
- Asynchronous Temporal Fields for Action Recognition
- Deep Reinforcement Learning for Wireless Sensor Scheduling in Cyber-Physical Systems
- Never Forget: Balancing Exploration and Exploitation via Learning Optical Flow
- Sequential Evaluation and Generation Framework for Combinatorial Recommender System
- Learning Runtime Parameters in Computer Systems with Delayed Experience Injection
- Convergence of Value Aggregation for Imitation Learning
- Why Pay More When You Can Pay Less: A Joint Learning Framework for Active Feature Acquisition and Classification
- Meta learning Framework for Automated Driving
- Improving Image Classification Robustness through Selective CNN-Filters Fine-Tuning
- PaintBot: A Reinforcement Learning Approach for Natural Media Painting
- ROS2Learn: a reinforcement learning framework for ROS 2
- Continuous-action Reinforcement Learning for Playing Racing Games: Comparing SPG to PPO
- Learning Sparse Representations Incrementally in Deep Reinforcement Learning
- Learned Hardware/Software Co-Design of Neural Accelerators
- DMRO:A Deep Meta Reinforcement Learning-based Task Offloading Framework for Edge-Cloud Computing
- An Intelligent Control Strategy for buck DC-DC Converter via Deep Reinforcement Learning
- Momentum Q-learning with Finite-Sample Convergence Guarantee
- Control-Aware Representations for Model-based Reinforcement Learning
- Learning to Track Dynamic Targets in Partially Known Environments
- FlapAI Bird: Training an Agent to Play Flappy Bird Using Reinforcement Learning Techniques
- Optimizing for the Future in Non-Stationary MDPs
- A deep learning theory for neural networks grounded in physics
- A review of motion planning algorithms for intelligent robotics
- Scalable Voltage Control using Structure-Driven Hierarchical Deep Reinforcement Learning
- Self-Imitation Advantage Learning
- Batch-Constrained Distributional Reinforcement Learning for Session-based Recommendation
- Fever Basketball: A Complex, Flexible, and Asynchronized Sports Game Environment for Multi-agent Reinforcement Learning
- Autonomous Blimp Control using Deep Reinforcement Learning
- The Benchmark Lottery
- A Deep Reinforcement Learning Approach for Traffic Signal Control Optimization
- Drone swarm patrolling with uneven coverage requirements
- Which Mutual-Information Representation Learning Objectives are Sufficient for Control?
- Deception in Social Learning: A Multi-Agent Reinforcement Learning Perspective
- TiKick: Towards Playing Multi-agent Football Full Games from Single-agent Demonstrations
- Towards Scalable Verification of Deep Reinforcement Learning
- Capability Iteration Network for Robot Path Planning
- Full Gradient DQN Reinforcement Learning: A Provably Convergent Scheme
- Multimodal Safety-Critical Scenarios Generation for Decision-Making Algorithms Evaluation
- SuperSuit: Simple Microwrappers for Reinforcement Learning Environments
- Delay-Aware Model-Based Reinforcement Learning for Continuous Control
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Leveraging End-to-End Speech Recognition with Neural Architecture Search
- Adaptive Online Planning for Continual Lifelong Learning
- Robust Opponent Modeling via Adversarial Ensemble Reinforcement Learning in Asymmetric Imperfect-Information Games
- Prediction, Consistency, Curvature: Representation Learning for Locally-Linear Control
- Regional Tree Regularization for Interpretability in Black Box Models
- Control of nonlinear, complex and black-boxed greenhouse system with reinforcement learning
- Risk-Sensitive Compact Decision Trees for Autonomous Execution in Presence of Simulated Market Response
- Generalized Second Order Value Iteration in Markov Decision Processes
- Neural-encoding Human Experts' Domain Knowledge to Warm Start Reinforcement Learning
- Real-Time Fine-Grained Air Quality Sensing Networks in Smart City: Design, Implementation and Optimization
- Automated Image Data Preprocessing with Deep Reinforcement Learning
- Learning Intelligent Dialogs for Bounding Box Annotation
- Learning Deep Mean Field Games for Modeling Large Population Behavior
- Learning Light Transport the Reinforced Way
- Using Meta Reinforcement Learning to Bridge the Gap between Simulation and Experiment in Energy Demand Response
- Joint Policy Search for Multi-agent Collaboration with Imperfect Information
- Is Plug-in Solver Sample-Efficient for Feature-based Reinforcement Learning?
- Planning Robot Motion using Deep Visual Prediction
- Deep Reinforcement Learning for Demand Driven Services in Logistics and Transportation Systems: A Survey
- A Framework for Studying Reinforcement Learning and Sim-to-Real in Robot Soccer
- Prototypical Q Networks for Automatic Conversational Diagnosis and Few-Shot New Disease Adaption
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey
- Learning hierarchical behavior and motion planning for autonomous driving
- Emergent Road Rules In Multi-Agent Driving Environments
- Influencing Towards Stable Multi-Agent Interactions
- SOAC: The Soft Option Actor-Critic Architecture
- Learning to Prune Deep Neural Networks via Reinforcement Learning
- Jelly Bean World: A Testbed for Never-Ending Learning
- Augmenting GAIL with BC for sample efficient imitation learning
- Learning to Sample the Most Useful Training Patches from Images
- Deep Reinforcement Learning with Embedded LQR Controllers
- Optimizing Routerless Network-on-Chip Designs: An Innovative Learning-Based Framework
- Combined Peak Reduction and Self-Consumption Using Proximal Policy Optimization
- How to Make Deep RL Work in Practice
- Cross-Domain Perceptual Reward Functions
- Finding the best design parameters for optical nanostructures using reinforcement learning
- Continuous Coordination As a Realistic Scenario for Lifelong Learning
- Sim-Real Joint Reinforcement Transfer for 3D Indoor Navigation
- QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning
- Dynamic Control of a Fiber Manufacturing Process using Deep Reinforcement Learning
- AUBER: Automated BERT Regularization
- Reinforcement Learning for Load-balanced Parallel Particle Tracing
- PixelRL: Fully Convolutional Network with Reinforcement Learning for Image Processing
- Contrastive Variational Reinforcement Learning for Complex Observations
- A Deep Q-learning/genetic Algorithms Based Novel Methodology For Optimizing Covid-19 Pandemic Government Actions
- Reinforcement Learning-Based Coverage Path Planning with Implicit Cellular Decomposition
- Hierarchical Reinforcement Learning Method for Autonomous Vehicle Behavior Planning
- Weighted Double Deep Multiagent Reinforcement Learning in Stochastic Cooperative Environments
- A Decentralized Policy Gradient Approach to Multi-task Reinforcement Learning
- Graph2Seq: Scalable Learning Dynamics for Graphs
- On Reinforcement Learning for Full-length Game of StarCraft
- Gradient-based Training of Slow Feature Analysis by Differentiable Approximate Whitening
- Meta Dialogue Policy Learning
- Data Driven Control with Learned Dynamics: Model-Based versus Model-Free Approach
- Improving Conditional Sequence Generative Adversarial Networks by Stepwise Evaluation
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
- AgentGraph: Towards Universal Dialogue Management with Structured Deep Reinforcement Learning
- Relative stability toward diffeomorphisms indicates performance in deep nets
- Maximum Resilience of Artificial Neural Networks
- Deep Reinforcement Learning in Quantitative Algorithmic Trading: A Review
- A Probabilistic Simulator of Spatial Demand for Product Allocation
- Reinforcement Learning for Fair Dynamic Pricing
- Implicit Policy for Reinforcement Learning
- When Simple Exploration is Sample Efficient: Identifying Sufficient Conditions for Random Exploration to Yield PAC RL Algorithms
- QoS-Aware Scheduling in New Radio Using Deep Reinforcement Learning
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement Learning
- Reinforcement Learning based Multi-Access Control and Battery Prediction with Energy Harvesting in IoT Systems
- X-ToM: Explaining with Theory-of-Mind for Gaining Justified Human Trust
- Rethinking Exposure Bias In Language Modeling
- An A* Curriculum Approach to Reinforcement Learning for RGBD Indoor Robot Navigation
- Plan2Vec: Unsupervised Representation Learning by Latent Plans
- Towards Lifelong Self-Supervision: A Deep Learning Direction for Robotics
- Learning Vision-based Robotic Manipulation Tasks Sequentially in Offline Reinforcement Learning Settings
- AI Enabling Technologies: A Survey
- Reinforcement Learning of Beam Codebooks in Millimeter Wave and Terahertz MIMO Systems
- Jointly-Learned State-Action Embedding for Efficient Reinforcement Learning
- Automatic Recall Machines: Internal Replay, Continual Learning and the Brain
- Policy Learning Using Weak Supervision
- Deep reinforcement learning to detect brain lesions on MRI: a proof-of-concept application of reinforcement learning to medical images
- Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous Time
- Toybox: A Suite of Environments for Experimental Evaluation of Deep Reinforcement Learning
- PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control
- Hindsight Expectation Maximization for Goal-conditioned Reinforcement Learning
- Effects of Loss Functions And Target Representations on Adversarial Robustness
- Idiosyncrasies and challenges of data driven learning in electronic trading
- Deep Reinforcement Learning-based UAV Navigation and Control: A Soft Actor-Critic with Hindsight Experience Replay Approach
- Real-time visual tracking by deep reinforced decision making
- A Function Approximation Method for Model-based High-Dimensional Inverse Reinforcement Learning
- Refactoring Neural Networks for Verification
- Simulation of machine learning-based 6G systems in virtual worlds
- Strategies for Conceptual Change in Convolutional Neural Networks
- Nearly Horizon-Free Offline Reinforcement Learning
- Sampled Policy Gradient for Learning to Play the Game Agar.io
- Towards continuous control of flippers for a multi-terrain robot using deep reinforcement learning
- Liquid Splash Modeling with Neural Networks
- The Effects of Memory Replay in Reinforcement Learning
- Unknowable Manipulators: Social Network Curator Algorithms
- Online Hyper-parameter Tuning in Off-policy Learning via Evolutionary Strategies
- Task-Agnostic Morphology Evolution
- Imitation Learning for Human Pose Prediction
- A Deep Reinforcement Learning Approach to Efficient Drone Mobility Support
- MoTiAC: Multi-Objective Actor-Critics for Real-Time Bidding
- A Visual Communication Map for Multi-Agent Deep Reinforcement Learning
- Computation on Sparse Neural Networks: an Inspiration for Future Hardware
- Resolving Spurious Correlations in Causal Models of Environments via Interventions
- Group Fairness in Bandit Arm Selection
- Efficient Deep Reinforcement Learning via Adaptive Policy Transfer
- Remote Sensing and Machine Learning for Food Crop Production Data in Africa Post-COVID-19
- Towards Safe Control of Continuum Manipulator Using Shielded Multiagent Reinforcement Learning
- Neural Packet Classification
- Combinational Q-Learning for Dou Di Zhu
- Action-Sufficient State Representation Learning for Control with Structural Constraints
- Towards Understanding the Generalization Bias of Two Layer Convolutional Linear Classifiers with Gradient Descent
- ToyBox: Better Atari Environments for Testing Reinforcement Learning Agents
- Trust Region Value Optimization using Kalman Filtering
- Microwave Integrated Circuits Design with Relational Induction Neural Network
- Balance Between Efficient and Effective Learning: Dense2Sparse Reward Shaping for Robot Manipulation with Environment Uncertainty
- Nonparametric Stochastic Compositional Gradient Descent for Q-Learning in Continuous Markov Decision Problems
- Reconstruction of training samples from loss functions
- Understanding Adversarial Attacks on Observations in Deep Reinforcement Learning
- Teaching Machines to Converse
- Deep Reinforcement Learning for Doom using Unsupervised Auxiliary Tasks
- ANT: Learning Accurate Network Throughput for Better Adaptive Video Streaming
- Malaria Likelihood Prediction By Effectively Surveying Households Using Deep Reinforcement Learning
- Generic Itemset Mining Based on Reinforcement Learning
- On architectural choices in deep learning: From network structure to gradient convergence and parameter estimation
- IOS: Inter-Operator Scheduler for CNN Acceleration
- Visual Concept Recognition and Localization via Iterative Introspection
- RBED: Reward Based Epsilon Decay
- Reinforcement Learning, Bit by Bit
- How can AI Automate End-to-End Data Science?
- Constrained Model-Free Reinforcement Learning for Process Optimization
- Transfer Learning versus Multi-agent Learning regarding Distributed Decision-Making in Highway Traffic
- Domain Adaptation In Reinforcement Learning Via Latent Unified State Representation
- Autonomous Driving using Safe Reinforcement Learning by Incorporating a Regret-based Human Lane-Changing Decision Model
- PsiPhi-Learning: Reinforcement Learning with Demonstrations using Successor Features and Inverse Temporal Difference Learning
- Model-free and Bayesian Ensembling Model-based Deep Reinforcement Learning for Particle Accelerator Control Demonstrated on the FERMI FEL
- Real-time Artificial Intelligence for Accelerator Control: A Study at the Fermilab Booster
- Adversarial Reinforcement Learning under Partial Observability in Autonomous Computer Network Defence
- Interactive Visualization for Debugging RL
- Dynamic Control of Stochastic Evolution: A Deep Reinforcement Learning Approach to Adaptively Targeting Emergent Drug Resistance
- Engineering problems in machine learning systems
- Deep RL With Information Constrained Policies: Generalization in Continuous Control
- Soft Expert Reward Learning for Vision-and-Language Navigation
- Active World Model Learning with Progress Curiosity
- Multiplayer Support for the Arcade Learning Environment
- GRAC: Self-Guided and Self-Regularized Actor-Critic
- Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
- MUSE: Modularizing Unsupervised Sense Embeddings
- Policy-GNN: Aggregation Optimization for Graph Neural Networks
- Learning "What-if" Explanations for Sequential Decision-Making
- Efficient Ridesharing Dispatch Using Multi-Agent Reinforcement Learning
- Implicit Distributional Reinforcement Learning
- Learning to Communicate Using Counterfactual Reasoning
- Developing a Simple Model for Sand-Tool Interaction and Autonomously Shaping Sand
- Incremental Reinforcement Learning --- a New Continuous Reinforcement Learning Frame Based on Stochastic Differential Equation methods
- Meta-Learning Bandit Policies by Gradient Ascent
- Solving Hard AI Planning Instances Using Curriculum-Driven Deep Reinforcement Learning
- Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning
- Attend, Adapt and Transfer: Attentive Deep Architecture for Adaptive Transfer from multiple sources in the same domain
- An FPGA-Based On-Device Reinforcement Learning Approach using Online Sequential Learning
- FIXAR: A Fixed-Point Deep Reinforcement Learning Platform with Quantization-Aware Training and Adaptive Parallelism
- Ultrasound-Guided Robotic Navigation with Deep Reinforcement Learning
- Intelligent Coordination among Multiple Traffic Intersections Using Multi-Agent Reinforcement Learning
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Policies
- Weighing Counts: Sequential Crowd Counting by Reinforcement Learning
- Multiagent Rollout and Policy Iteration for POMDP with Application to Multi-Robot Repair Problems
- Learning to Design Games: Strategic Environments in Reinforcement Learning
- A Game-Theoretic Approach to Multi-Agent Trust Region Optimization
- Towards Trainable Media: Using Waves for Neural Network-Style Training
- Self-Tuning Sectorization: Deep Reinforcement Learning Meets Broadcast Beam Optimization
- Model Based Planning with Energy Based Models
- Improving a sequence-to-sequence nlp model using a reinforcement learning policy algorithm
- TFPnP: Tuning-free Plug-and-Play Proximal Algorithm with Applications to Inverse Imaging Problems
- Obstacle Avoidance and Navigation Utilizing Reinforcement Learning with Reward Shaping
- Deep Reinforcement Learning for Adaptive Learning Systems
- Towards White-box Benchmarks for Algorithm Control
- Prophet: Proactive Candidate-Selection for Federated Learning by Predicting the Qualities of Training and Reporting Phases
- Collaborative Policy Learning for Open Knowledge Graph Reasoning
- Towards Optimal District Heating Temperature Control in China with Deep Reinforcement Learning
- Implicit Generative Modeling for Efficient Exploration
- Evaluating the Rainbow DQN Agent in Hanabi with Unseen Partners
- LPaintB: Learning to Paint from Self-Supervision
- A Deep Reinforcement Learning Based Multi-Criteria Decision Support System for Textile Manufacturing Process Optimization
- Off-Policy Adversarial Inverse Reinforcement Learning
- Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction
- Deep Reinforcement Learning for Optimal Stopping with Application in Financial Engineering
- Efficient and Effective Similar Subtrajectory Search with Deep Reinforcement Learning
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Offline Reinforcement Learning Hands-On
- Deep Learning with Experience Ranking Convolutional Neural Network for Robot Manipulator
- Linear Representation Meta-Reinforcement Learning for Instant Adaptation
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- Learning to Score Behaviors for Guided Policy Optimization
- Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
- Smart Train Operation Algorithms based on Expert Knowledge and Reinforcement Learning
- TauRieL: Targeting Traveling Salesman Problem with a deep reinforcement learning inspired architecture
- Deep Reinforcement Learning-Based Product Recommender for Online Advertising
- Dampen the Stop-and-Go Traffic with Connected and Automated Vehicles -- A Deep Reinforcement Learning Approach
- No-Regret Reinforcement Learning with Heavy-Tailed Rewards
- Synthesis of Discounted-Reward Optimal Policies for Markov Decision Processes Under Linear Temporal Logic Specifications
- Combining Reinforcement Learning with Model Predictive Control for On-Ramp Merging
- Watch the Unobserved: A Simple Approach to Parallelizing Monte Carlo Tree Search
- Taylor Expansion Policy Optimization
- Deep Reinforcement Learning for Producing Furniture Layout in Indoor Scenes
- Low Dimensional State Representation Learning with Reward-shaped Priors
- Off-Policy Actor-Critic with Shared Experience Replay
- AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control
- Simultaneous Decision Making for Stochastic Multi-echelon Inventory Optimization with Deep Neural Networks as Decision Makers
- A General Approach for Using Deep Neural Network for Digital Watermarking
- Corner Case Generation and Analysis for Safety Assessment of Autonomous Vehicles
- Testing match-3 video games with Deep Reinforcement Learning
- Packet Routing with Graph Attention Multi-agent Reinforcement Learning
- Deep Reinforcement Learning via L-BFGS Optimization
- Deep Stock Trading: A Hierarchical Reinforcement Learning Framework for Portfolio Optimization and Order Execution
- Interactive Language Acquisition with One-shot Visual Concept Learning through a Conversational Game
- Out-of-Distribution Dynamics Detection: RL-Relevant Benchmarks and Results
- Automated Gain Control Through Deep Reinforcement Learning for Downstream Radar Object Detection
- The Adversarial Resilience Learning Architecture for AI-based Modelling, Exploration, and Operation of Complex Cyber-Physical Systems
- Quantum Neural Networks: Concepts, Applications, and Challenges
- Learning Graph Structure With A Finite-State Automaton Layer
- On Deep Representation Learning from Noisy Web Images
- High-level Features for Resource Economy and Fast Learning in Skill Transfer
- Distributional Reinforcement Learning for Multi-Dimensional Reward Functions
- Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
- Comprehensive and Efficient Data Labeling via Adaptive Model Scheduling
- Towards General Function Approximation in Zero-Sum Markov Games
- Multi-Agent Path Planning based on MPC and DDPG
- Predictive Coding for Locally-Linear Control
- A Stochastic Game Framework for Efficient Energy Management in Microgrid Networks
- Optimizing Taxi Carpool Policies via Reinforcement Learning and Spatio-Temporal Mining
- A Comprehensive Survey on the Ambulance Routing and Location Problems
- AutoSoC: Automating Algorithm-SOC Co-design for Aerial Robots
- Estimating Q(s,s') with Deep Deterministic Dynamics Gradients
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional Smoothing
- On-Policy Trust Region Policy Optimisation with Replay Buffers
- Learning from Learning Machines: Optimisation, Rules, and Social Norms
- Planning with a Receding Horizon for Manipulation in Clutter using a Learned Value Function
- Convex Q-Learning, Part 1: Deterministic Optimal Control
- A Statistical Theory of Deep Learning via Proximal Splitting
- When Multiple Agents Learn to Schedule: A Distributed Radio Resource Management Framework
- Coordination in Adversarial Sequential Team Games via Multi-Agent Deep Reinforcement Learning
- Are Gradient-based Saliency Maps Useful in Deep Reinforcement Learning?
- Neural Architecture Evolution in Deep Reinforcement Learning for Continuous Control
- Multi-Objective Optimization of the Textile Manufacturing Process Using Deep-Q-Network Based Multi-Agent Reinforcement Learning
- Driving Tasks Transfer in Deep Reinforcement Learning for Decision-making of Autonomous Vehicles
- Continuous Homeostatic Reinforcement Learning for Self-Regulated Autonomous Agents
- P3O: Policy-on Policy-off Policy Optimization
- Transfer Learning Across Patient Variations with Hidden Parameter Markov Decision Processes
- Towards Personalized Dialog Policies for Conversational Skill Discovery
- Variance Reduction for Deep Q-Learning using Stochastic Recursive Gradient
- Unsupervised Visual Attention and Invariance for Reinforcement Learning
- MRPB 1.0: A Unified Benchmark for the Evaluation of Mobile Robot Local Planning Approaches
- Synthesizing Chemical Plant Operation Procedures using Knowledge, Dynamic Simulation and Deep Reinforcement Learning
- Iterative Amortized Policy Optimization
- Measuring Human Adaptation to AI in Decision Making: Application to Evaluate Changes after AlphaGo
- Join Query Optimization with Deep Reinforcement Learning Algorithms
- Symphony from Synapses: Neocortex as a Universal Dynamical Systems Modeller using Hierarchical Temporal Memory
- Competitive Experience Replay
- A Visual Embedding for the Unsupervised Extraction of Abstract Semantics
- Learning Representations in Reinforcement Learning:An Information Bottleneck Approach
- BRAC+: Improved Behavior Regularized Actor Critic for Offline Reinforcement Learning
- On Instrumental Variable Regression for Deep Offline Policy Evaluation
- Coordinated Heterogeneous Distributed Perception based on Latent Space Representation
- Acquiring Target Stacking Skills by Goal-Parameterized Deep Reinforcement Learning
- A Crash Course on Reinforcement Learning
- Continual Learning via Bit-Level Information Preserving
- CoordiQ : Coordinated Q-learning for Electric Vehicle Charging Recommendation
- Towards Understanding Chinese Checkers with Heuristics, Monte Carlo Tree Search, and Deep Reinforcement Learning
- Iterative Refinement of the Approximate Posterior for Directed Belief Networks
- Reinforced Imitation Learning by Free Energy Principle
- Task-oriented Design through Deep Reinforcement Learning
- Low Precision Policy Distillation with Application to Low-Power, Real-time Sensation-Cognition-Action Loop with Neuromorphic Computing
- Visual Diagnostics for Deep Reinforcement Learning Policy Development
- Personalized Multimorbidity Management for Patients with Type 2 Diabetes Using Reinforcement Learning of Electronic Health Records
- Learning Automata Based Q-learning for Content Placement in Cooperative Caching
- Reinforcement Learning for Intelligent Healthcare Systems: A Comprehensive Survey
- Bayesian Optimization for Iterative Learning
- Automatic Data Augmentation by Learning the Deterministic Policy
- Reinforcement Learning with Efficient Active Feature Acquisition
- Learning Manipulation under Physics Constraints with Visual Perception
- Deep Q-Network for Angry Birds
- Using Reinforcement Learning to Validate Empirical Game-Theoretic Analysis: A Continuous Double Auction Study
- Evolving Neural Networks in Reinforcement Learning by means of UMDAc
- PreCo: A Large-scale Dataset in Preschool Vocabulary for Coreference Resolution
- Curiosity-Driven Recommendation Strategy for Adaptive Learning via Deep Reinforcement Learning
- Borrowing From the Future: Addressing Double Sampling in Model-free Control
- AutoPilot: Automating SoC Design Space Exploration for SWaP Constrained Autonomous UAVs
- Intrinsic Geometric Vulnerability of High-Dimensional Artificial Intelligence
- Sim-To-Real Transfer for Miniature Autonomous Car Racing
- Reinforcement Learning based Dynamic Model Selection for Short-Term Load Forecasting
- Delay Constrained Buffer-Aided Relay Selection in the Internet of Things with Decision-Assisted Reinforcement Learning
- Information Bottleneck in Control Tasks with Recurrent Spiking Neural Networks
- Playing Flappy Bird via Asynchronous Advantage Actor Critic Algorithm
- Solving The Lunar Lander Problem under Uncertainty using Reinforcement Learning
- Reward-estimation variance elimination in sequential decision processes
- Co-training for Policy Learning
- Fairness-Oriented User Scheduling for Bursty Downlink Transmission Using Multi-Agent Reinforcement Learning
- Tensor-based Cooperative Control for Large Scale Multi-intersection Traffic Signal Using Deep Reinforcement Learning and Imitation Learning
- DISPATCH: Design Space Exploration of Cyber-Physical Systems
- ADARES: Adaptive Resource Management for Virtual Machines
- Zero-Shot Learning of Text Adventure Games with Sentence-Level Semantics
- Dynamic Dispatching for Large-Scale Heterogeneous Fleet via Multi-agent Deep Reinforcement Learning
- Learning Agents With Prioritization and Parameter Noise in Continuous State and Action Space
- On Catastrophic Interference in Atari 2600 Games
- Optimizing Large-Scale Fleet Management on a Road Network using Multi-Agent Deep Reinforcement Learning with Graph Neural Network
- Deep Echo State Q-Network (DEQN) and Its Application in Dynamic Spectrum Sharing for 5G and Beyond
- An Energy-Saving Snake Locomotion Gait Policy Obtained Using Deep Reinforcement Learning
- AutoEG: Automated Experience Grafting for Off-Policy Deep Reinforcement Learning
- MobileVisFixer: Tailoring Web Visualizations for Mobile Phones Leveraging an Explainable Reinforcement Learning Framework
- Financial Crime & Fraud Detection Using Graph Computing: Application Considerations & Outlook
- OER: Offline Experience Replay for Continual Offline Reinforcement Learning
- Competitive Multi-Agent Deep Reinforcement Learning with Counterfactual Thinking
- Multi-Agent Coordination in Adversarial Environments through Signal Mediated Strategies
- Distributed Machine Learning for Wireless Communication Networks: Techniques, Architectures, and Applications
- Disentangling causal effects for hierarchical reinforcement learning
- Proximal Policy Optimization Smoothed Algorithm
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- AdaDeep: A Usage-Driven, Automated Deep Model Compression Framework for Enabling Ubiquitous Intelligent Mobiles
- Learning Value Functions in Deep Policy Gradients using Residual Variance
- Correcting Experience Replay for Multi-Agent Communication
- The Architectural Implications of Distributed Reinforcement Learning on CPU-GPU Systems
- MELD: Meta-Reinforcement Learning from Images via Latent State Models
- Deep Reinforcement Learning based Dynamic Optimization of Bus Timetable
- Multi-Graph Tensor Networks
- Nonlinear Regression with a Convolutional Encoder-Decoder for Remote Monitoring of Surface Electrocardiograms
- Transferable Active Grasping and Real Embodied Dataset
- Temporal-difference learning with nonlinear function approximation: lazy training and mean field regimes
- Delayed Q-update: A novel credit assignment technique for deriving an optimal operation policy for the Grid-Connected Microgrid
- A Model-free Learning Algorithm for Infinite-horizon Average-reward MDPs with Near-optimal Regret
- Model-agnostic and Scalable Counterfactual Explanations via Reinforcement Learning
- Finding Needles in a Moving Haystack: Prioritizing Alerts with Adversarial Reinforcement Learning
- Blind interactive learning of modulation schemes: Multi-agent cooperation without co-design
- PBCS : Efficient Exploration and Exploitation Using a Synergy between Reinforcement Learning and Motion Planning
- Variational Empowerment as Representation Learning for Goal-Based Reinforcement Learning
- Network Automatic Pruning: Start NAP and Take a Nap
- Learning dynamic polynomial proofs
- Reinforcement Mechanism Design for e-commerce
- Player-AI Interaction: What Neural Network Games Reveal About AI as Play
- Reinforcement Learning for Flexibility Design Problems
- Supervised Learning and Reinforcement Learning of Feedback Models for Reactive Behaviors: Tactile Feedback Testbed
- DDPG++: Striving for Simplicity in Continuous-control Off-Policy Reinforcement Learning
- Adaptive Procedural Task Generation for Hard-Exploration Problems
- Continual Learning: Tackling Catastrophic Forgetting in Deep Neural Networks with Replay Processes
- Pseudo-Model-Free Hedging for Variable Annuities via Deep Reinforcement Learning
- Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRL
- Smooth Q-learning: Accelerate Convergence of Q-learning Using Similarity
- TanksWorld: A Multi-Agent Environment for AI Safety Research
- Disturbing Reinforcement Learning Agents with Corrupted Rewards
- Graph-based State Representation for Deep Reinforcement Learning
- Sample Efficient Feature Selection for Factored MDPs
- Policy Smoothing for Provably Robust Reinforcement Learning
- Optimal -Coverage Charging Problem
- Q-Networks for Binary Vector Actions
- Learning to be Global Optimizer
- Learning to Gather without Communication
- Adversarial recovery of agent rewards from latent spaces of the limit order book
- From Gameplay to Symbolic Reasoning: Learning SAT Solver Heuristics in the Style of Alpha(Go) Zero
- Hierarchically Integrated Models: Learning to Navigate from Heterogeneous Robots
- Discovering Latent States for Model Learning: Applying Sensorimotor Contingencies Theory and Predictive Processing to Model Context
- CostNet: An End-to-End Framework for Goal-Directed Reinforcement Learning
- Learning to solve arithmetic problems with a virtual abacus
- On Pathologies in KL-Regularized Reinforcement Learning from Expert Demonstrations
- Offline-Online Reinforcement Learning for Energy Pricing in Office Demand Response: Lowering Energy and Data Costs
- Learn Zero-Constraint-Violation Policy in Model-Free Constrained Reinforcement Learning
- Cooperative multi-agent reinforcement learning for high-dimensional nonequilibrium control
- A Free Lunch from the Noise: Provable and Practical Exploration for Representation Learning
- All-In-One: Artificial Association Neural Networks
- Learning Efficient Multi-Agent Cooperative Visual Exploration
- Double Deep Q-learning Based Real-Time Optimization Strategy for Microgrids
- A Simple Approach to Continual Learning by Transferring Skill Parameters
- Feedback Linearization of Car Dynamics for Racing via Reinforcement Learning
- Aligning an optical interferometer with beam divergence control and continuous action space
- Online Sub-Sampling for Reinforcement Learning with General Function Approximation
- DECORE: Deep Compression with Reinforcement Learning
- Non-decreasing Quantile Function Network with Efficient Exploration for Distributional Reinforcement Learning
- Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider Utilities
- Temporal-Difference Value Estimation via Uncertainty-Guided Soft Updates
- Learning to Communicate with Reinforcement Learning for an Adaptive Traffic Control System
- Improved Exploring Starts by Kernel Density Estimation-Based State-Space Coverage Acceleration in Reinforcement Learning
- Gym-RTS: Toward Affordable Full Game Real-time Strategy Games Research with Deep Reinforcement Learning
- Decentralized Multi-Agent Reinforcement Learning: An Off-Policy Method
- Touch-based Curiosity for Sparse-Reward Tasks
- Deep Reinforcement Learning From Raw Pixels in Doom
- Reinforcement Learning for Autonomous Driving with Latent State Inference and Spatial-Temporal Relationships
- Hedging of Financial Derivative Contracts via Monte Carlo Tree Search
- Improving Classifier Confidence using Lossy Label-Invariant Transformations
- Explaining Conditions for Reinforcement Learning Behaviors from Real and Imagined Data
- Learning Control Barrier Functions with High Relative Degree for Safety-Critical Control
- Solving optimal stopping problems with Deep Q-Learning
- Revocable Deep Reinforcement Learning with Affinity Regularization for Outlier-Robust Graph Matching
- Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation
- Analyzing Finite Neural Networks: Can We Trust Neural Tangent Kernel Theory?
- Exploring Fluent Query Reformulations with Text-to-Text Transformers and Reinforcement Learning
- Energy Aware Deep Reinforcement Learning Scheduling for Sensors Correlated in Time and Space
- Evaluating the Safety of Deep Reinforcement Learning Models using Semi-Formal Verification
- Multi-Agent Reinforcement Learning in NOMA-aided UAV Networks for Cellular Offloading
- A Dynamic Penalty Function Approach for Constraints-Handling in Reinforcement Learning
- Affordance as general value function: A computational model
- A Deep Reinforcement Learning Approach for Ramp Metering Based on Traffic Video Data
- OPAC: Opportunistic Actor-Critic
- Energy-based Surprise Minimization for Multi-Agent Value Factorization
- A Comparative Analysis of Deep Reinforcement Learning-enabled Freeway Decision-making for Automated Vehicles
- Active Screening for Recurrent Diseases: A Reinforcement Learning Approach
- Learning Complex Multi-Agent Policies in Presence of an Adversary
- Prediction and Generalisation over Directed Actions by Grid Cells
- Multi-Agent Reinforcement Learning of 3D Furniture Layout Simulation in Indoor Graphics Scenes
- An Interaction-aware Evaluation Method for Highly Automated Vehicles
- On the Ethics of Building AI in a Responsible Manner
- Improved Regret Bound and Experience Replay in Regularized Policy Iteration
- Generating Socially Acceptable Perturbations for Efficient Evaluation of Autonomous Vehicles
- Graph Computing for Financial Crime and Fraud Detection: Trends, Challenges and Outlook
- Using AI for Mitigating the Impact of Network Delay in Cloud-based Intelligent Traffic Signal Control
- Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
- Tensor Networks for Multi-Modal Non-Euclidean Data
- Learning medical triage from clinicians using Deep Q-Learning
- A coevolutionary approach to deep multi-agent reinforcement learning
- CropGym: a Reinforcement Learning Environment for Crop Management
- Context-Aware Deep Q-Network for Decentralized Cooperative Reconnaissance by a Robotic Swarm
- Towards Runtime Verification of Programmable Switches
- Training an Interactive Humanoid Robot Using Multimodal Deep Reinforcement Learning
- Deep Reinforcement Learning Based Spectrum Allocation in Integrated Access and Backhaul Networks
- Sample-Efficient Model-based Actor-Critic for an Interactive Dialogue Task
- Formal Methods with a Touch of Magic
- Alleviating Privacy Attacks via Causal Learning
- Smaller Models, Better Generalization
- Decorrelated Double Q-learning
- Progressive Relation Learning for Group Activity Recognition
- Optimistic Distributionally Robust Policy Optimization
- AiAds: Automated and Intelligent Advertising System for Sponsored Search
- Maximum Entropy Model Rollouts: Fast Model Based Policy Optimization without Compounding Errors
- Concept and the implementation of a tool to convert industry 4.0 environments modeled as FSM to an OpenAI Gym wrapper
- Expert Level control of Ramp Metering based on Multi-task Deep Reinforcement Learning
- Group Equivariant Deep Reinforcement Learning
- TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning?
- Simulating multi-exit evacuation using deep reinforcement learning
- Generalization to Novel Objects using Prior Relational Knowledge
- Beyond Exponentially Discounted Sum: Automatic Learning of Return Function
- Improved Hard Example Mining by Discovering Attribute-based Hard Person Identity
- Understanding how T helper cells learn to coordinate effective immune responses through the lens of reinforcement learning
- An Optimistic Acceleration of AMSGrad for Nonconvex Optimization
- Model-Free Episodic Control with State Aggregation
- Competitiveness of MAP-Elites against Proximal Policy Optimization on locomotion tasks in deterministic simulations
- Interactive Reinforcement Learning for Feature Selection with Decision Tree in the Loop
- Unsupervised Emergence of Spatial Structure from Sensorimotor Prediction
- Content-Aware Personalised Rate Adaptation for Adaptive Streaming via Deep Video Analysis
- Decentralized Multi-Agent Actor-Critic with Generative Inference
- Neural Program Synthesis By Self-Learning
- Coordination of PV Smart Inverters Using Deep Reinforcement Learning for Grid Voltage Regulation
- Towards a Reinforcement Learning Environment Toolbox for Intelligent Electric Motor Control
- Using reinforcement learning to minimize taxi idle times
- Affine Self Convolution
- Biologically inspired architectures for sample-efficient deep reinforcement learning
- NavigationNet: A Large-scale Interactive Indoor Navigation Dataset
- Reinforcement Learning for Autonomous Defence in Software-Defined Networking
- Deep RTS: A Game Environment for Deep Reinforcement Learning in Real-Time Strategy Games
- Fully Bayesian Recurrent Neural Networks for Safe Reinforcement Learning
- H2O-Cloud: A Resource and Quality of Service-Aware Task Scheduling Framework for Warehouse-Scale Data Centers -- A Hierarchical Hybrid DRL (Deep Reinforcement Learning) based Approach
- Saccadic Predictive Vision Model with a Fovea
- Robotic Grasp Manipulation Using Evolutionary Computing and Deep Reinforcement Learning
- Large Batch Training Does Not Need Warmup
- Cell Selection with Deep Reinforcement Learning in Sparse Mobile Crowdsensing
- Fast reinforcement learning for decentralized MAC optimization
- Learning Pregrasp Manipulation of Objects from Ungraspable Poses
- Intelligent Trainer for Model-Based Reinforcement Learning
- Informative Path Planning for Mobile Sensing with Reinforcement Learning
- Fractal AI: A fragile theory of intelligence
- Optimizing Deep Neural Networks with Multiple Search Neuroevolution
- Amplifying the Imitation Effect for Reinforcement Learning of UCAV's Mission Execution
- Towards Physically Safe Reinforcement Learning under Supervision
- Using Deep Q-learning To Prolong the Lifetime of Correlated Internet of Things Devices
- Typed Graph Networks
- Micro-Objective Learning : Accelerating Deep Reinforcement Learning through the Discovery of Continuous Subgoals
- Compression and Localization in Reinforcement Learning for ATARI Games
- Stochastic Lipschitz Q-Learning
- Recurrent Value Functions
- Deep Learning-Based Decoding of Constrained Sequence Codes
- Machine Learned Learning Machines
- Continual Reinforcement Learning with Diversity Exploration and Adversarial Self-Correction
- On the notion of number in humans and machines
- DeepNav: Learning to Navigate Large Cities
- Neuron ranking -- an informed way to condense convolutional neural networks architecture
- Iterative temporal differencing with random synaptic feedback weights support error backpropagation for deep learning
- Combinatorial Keyword Recommendations for Sponsored Search with Deep Reinforcement Learning
- DeepConfig: Automating Data Center Network Topologies Management with Machine Learning
- Boosting Offline Reinforcement Learning with Residual Generative Modeling
- Training like Playing: A Reinforcement Learning And Knowledge Graph-based framework for building Automatic Consultation System in Medical Field
- An Empirical Analysis of Proximal Policy Optimization with Kronecker-factored Natural Gradients
- Learning on Abstract Domains: A New Approach for Verifiable Guarantee in Reinforcement Learning
- Brittle AI, Causal Confusion, and Bad Mental Models: Challenges and Successes in the XAI Program
- Reinforcement Learning for Industrial Control Network Cyber Security Orchestration
- Metis: Multi-Agent Based Crisis Simulation System
- A New Concept of Deep Reinforcement Learning based Augmented General Sequence Tagging System
- Hybrid Policy Learning for Energy-Latency Tradeoff in MEC-Assisted VR Video Service
- Task-driven Semantic Coding via Reinforcement Learning
- Is Q-Learning Provably Efficient? An Extended Analysis
- A survey of benchmarking frameworks for reinforcement learning
- A Novel Deep Reinforcement Learning Based Stock Direction Prediction using Knowledge Graph and Community Aware Sentiments
- Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
- Soft Actor-Critic With Integer Actions
- Multiagent Reinforcement Learning based Energy Beamforming Control
- Slipping to the Extreme: A Mixed Method to Explain How Extreme Opinions Infiltrate Online Discussions
- Information theoretic analysis of computational models as a tool to understand the neural basis of behaviors
- Offline RL With Resource Constrained Online Deployment
- Comparing Heuristics, Constraint Optimization, and Reinforcement Learning for an Industrial 2D Packing Problem
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
- Multi-agent Policy Optimization with Approximatively Synchronous Advantage Estimation
- Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version
- Pareto Deterministic Policy Gradients and Its Application in 5G Massive MIMO Networks
- Low-Dimensional State and Action Representation Learning with MDP Homomorphism Metrics
- GAPLE: Generalizable Approaching Policy LEarning for Robotic Object Searching in Indoor Environment
- Knowledge Grounded Conversational Symptom Detection with Graph Memory Networks
- Off-policy Reinforcement Learning with Optimistic Exploration and Distribution Correction
- Fault-Tolerant Federated Reinforcement Learning with Theoretical Guarantee
- Learning Pessimism for Robust and Efficient Off-Policy Reinforcement Learning
- Temporal Abstraction in Reinforcement Learning with the Successor Representation
- Neural Network Based Nonlinear Weighted Finite Automata
- Learning Principle of Least Action with Reinforcement Learning
- Frequency Pooling: Shift-Equivalent and Anti-Aliasing Downsampling
- Explainable Biomedical Recommendations via Reinforcement Learning Reasoning on Knowledge Graphs
- Learning Transition Models with Time-delayed Causal Relations
- Convolutional Reservoir Computing for World Models
- Distributed off-Policy Actor-Critic Reinforcement Learning with Policy Consensus
- Distributed Multi-Agent Deep Reinforcement Learning Framework for Whole-building HVAC Control
- Deep Hierarchical Reinforcement Learning Based Recommendations via Multi-goals Abstraction
- An End-to-End Robot Architecture to Manipulate Non-Physical State Changes of Objects
- Representation based and Attention augmented Meta learning
- Dual Behavior Regularized Reinforcement Learning
- Policy Gradient for Continuing Tasks in Non-stationary Markov Decision Processes
- Instance-Aware Predictive Navigation in Multi-Agent Environments
- Improved Learning in Evolution Strategies via Sparser Inter-Agent Network Topologies
- Stabilizing Transformer-Based Action Sequence Generation For Q-Learning
- Pseudorehearsal in value function approximation
- Autonomy 2.0: Why is self-driving always 5 years away?
- Vizarel: A System to Help Better Understand RL Agents
- Practical Convex Formulation of Robust One-hidden-layer Neural Network Training
- Improving Experience Replay through Modeling of Similar Transitions' Sets
- Visual Analogies between Atari Games for Studying Transfer Learning in RL
- Deep Q learning for fooling neural networks
- Integrating Motion into Vision Models for Better Visual Prediction
- Reinforcement Learning with Convolutional Reservoir Computing
- Act to Reason: A Dynamic Game Theoretical Model of Driving
- Reducing Catastrophic Forgetting in Modular Neural Networks by Dynamic Information Balancing
- Hierarchical Reinforcement Learning Framework towards Multi-agent Navigation
- Assessing and Accelerating Coverage in Deep Reinforcement Learning
- Self-Calibrating Active Binocular Vision via Active Efficient Coding with Deep Autoencoders
- Pessimistic Model Selection for Offline Deep Reinforcement Learning
- Quantum Machine Learning For Classical Data
- Value-Based Reinforcement Learning for Continuous Control Robotic Manipulation in Multi-Task Sparse Reward Settings
- Chrome Dino Run using Reinforcement Learning
- InfoRL: Interpretable Reinforcement Learning using Information Maximization
- MMD-MIX: Value Function Factorisation with Maximum Mean Discrepancy for Cooperative Multi-Agent Reinforcement Learning
- Goal Reasoning by Selecting Subgoals with Deep Q-Learning
- Don't Forget Your Teacher: A Corrective Reinforcement Learning Framework
- Decision-Making in Reinforcement Learning
- Meta Arcade: A Configurable Environment Suite for Meta-Learning
- Neural MMO v1.3: A Massively Multiagent Game Environment for Training and Evaluating Neural Networks
- Detecting Adversarial Samples Using Density Ratio Estimates
- Learning active learning at the crossroads? evaluation and discussion
- Decentralized Multi-Agents by Imitation of a Centralized Controller
- Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms
- Amanuensis: The Programmer's Apprentice
- Deterministic Policy Gradients With General State Transitions
- Task-Relevant Object Discovery and Categorization for Playing First-person Shooter Games
- Improving width-based planning with compact policies
- On the Sensory Commutativity of Action Sequences for Embodied Agents
- Physical Reasoning Using Dynamics-Aware Models
- NVCell: Standard Cell Layout in Advanced Technology Nodes with Reinforcement Learning
- Developing Robust Digital Twins and Reinforcement Learning for Accelerator Control Systems at the Fermilab Booster
- Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling
- Learning task-agnostic representation via toddler-inspired learning
- IQ-Learn: Inverse soft-Q Learning for Imitation
- Collaborative Deep Reinforcement Learning for Joint Object Search
- Behavior Planning at Urban Intersections through Hierarchical Reinforcement Learning
- Minimalistic Attacks: How Little it Takes to Fool a Deep Reinforcement Learning Policy
- Deictic Image Maps: An Abstraction For Learning Pose Invariant Manipulation Policies
- NeoNav: Improving the Generalization of Visual Navigation via Generating Next Expected Observations
- Efficient and Phase-aware Video Super-resolution for Cardiac MRI
- Safe Reinforcement Learning for Grid Voltage Control
- Neuro-evolutionary Frameworks for Generalized Learning Agents
- Policy learning in SE(3) action spaces
- Discrete-to-Deep Supervised Policy Learning
- Active Perception in Adversarial Scenarios using Maximum Entropy Deep Reinforcement Learning
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Fast Retinomorphic Event Stream for Video Recognition and Reinforcement Learning
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy Evaluation
- Image Synthesis for Data Augmentation in Medical CT using Deep Reinforcement Learning
- Fast Online Exact Solutions for Deterministic MDPs with Sparse Rewards
- Correcting Momentum in Temporal Difference Learning
- Unsupervised Learning of Solutions to Differential Equations with Generative Adversarial Networks
- Data Efficient Training for Reinforcement Learning with Adaptive Behavior Policy Sharing
- Comparing heterogeneous entities using artificial neural networks of trainable weighted structural components and machine-learned activation functions
- FLASH: Fast Neural Architecture Search with Hardware Optimization
- Reinforcement Learning and Video Games
- A Gradient Estimator for Time-Varying Electrical Networks with Non-Linear Dissipation
- On Hyper-parameter Tuning for Stochastic Optimization Algorithms
- Self-Net: Lifelong Learning via Continual Self-Modeling
- Instance based Generalization in Reinforcement Learning
- Goal-constrained Sparse Reinforcement Learning for End-to-End Driving
- Designing Interpretable Approximations to Deep Reinforcement Learning
- Modular Object-Oriented Games: A Task Framework for Reinforcement Learning, Psychology, and Neuroscience
- Momentum-based Accelerated Q-learning
- Neuro-Symbolic Reinforcement Learning with First-Order Logic
- Improve Agents without Retraining: Parallel Tree Search with Off-Policy Correction
- C-3PO: Cyclic-Three-Phase Optimization for Human-Robot Motion Retargeting based on Reinforcement Learning
- Generative Question Refinement with Deep Reinforcement Learning in Retrieval-based QA System
- Controlling an Inverted Pendulum with Policy Gradient Methods-A Tutorial
- Learning Socially Appropriate Robot Approaching Behavior Toward Groups using Deep Reinforcement Learning
- Deep Learning: Our Miraculous Year 1990-1991
- Dirichlet Pruning for Neural Network Compression
- Multi-Issue Bargaining With Deep Reinforcement Learning
- Why? Why not? When? Visual Explanations of Agent Behavior in Reinforcement Learning
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Momentum Centering and Asynchronous Update for Adaptive Gradient Methods
- Robusta: Robust AutoML for Feature Selection via Reinforcement Learning
- Benchmarking Deep Graph Generative Models for Optimizing New Drug Molecules for COVID-19
- An overall view of key problems in algorithmic trading and recent progress
- Discrete Action On-Policy Learning with Action-Value Critic
- Physics-informed Dyna-Style Model-Based Deep Reinforcement Learning for Dynamic Control
- Grounding Hierarchical Reinforcement Learning Models for Knowledge Transfer
- An advantage actor-critic algorithm for robotic motion planning in dense and dynamic scenarios
- Give me a hint! Navigating Image Databases using Human-in-the-loop Feedback
- Deep Reinforcement Learning-based Task Offloading in Satellite-Terrestrial Edge Computing Networks
- Auditing Robot Learning for Safety and Compliance during Deployment
- D-ACC: Dynamic Adaptive Cruise Control for Highways with Ramps Based on Deep Q-Learning
- Maliva: Using Machine Learning to Rewrite Visualization Queries Under Time Constraints
- Feudal Steering: Hierarchical Learning for Steering Angle Prediction
- Recurrent Sum-Product-Max Networks for Decision Making in Perfectly-Observed Environments
- Interactive Machine Comprehension with Dynamic Knowledge Graphs
- Integrating Deep Learning and Augmented Reality to Enhance Situational Awareness in Firefighting Environments
- POAR: Efficient Policy Optimization via Online Abstract State Representation Learning
- Parallel Actors and Learners: A Framework for Generating Scalable RL Implementations
- Where Do Human Heuristics Come From?
- Bridging Cognitive Programs and Machine Learning
- Universal Memory Architectures for Autonomous Machines
- Training Transition Policies via Distribution Matching for Complex Tasks
- Realizing Continual Learning through Modeling a Learning System as a Fiber Bundle
- A study of first-passage time minimization via Q-learning in heated gridworlds
- Gap-Dependent Bounds for Two-Player Markov Games
- Towards Learning Generalizable Driving Policies from Restricted Latent Representations
- Augment-Reinforce-Merge Policy Gradient for Binary Stochastic Policy
- Towards Learning to Speak and Hear Through Multi-Agent Communication over a Continuous Acoustic Channel
- Integral Equations and Machine Learning
- How Crucial Is It for 6G Networks to Be Autonomous?
- Towards Brain-inspired System: Deep Recurrent Reinforcement Learning for Simulated Self-driving Agent
- Successor Feature Neural Episodic Control
- CubeTR: Learning to Solve The Rubiks Cube Using Transformers
- Human-Guided Learning of Column Networks: Augmenting Deep Learning with Advice
- Disentangling Options with Hellinger Distance Regularizer
- Overcoming Digital Gravity when using AI in Public Health Decisions
- Monte Carlo Tree Search for high precision manufacturing
- DriverGym: Democratising Reinforcement Learning for Autonomous Driving
- A Broad-persistent Advising Approach for Deep Interactive Reinforcement Learning in Robotic Environments
- Error Controlled Actor-Critic
- Object Exchangeability in Reinforcement Learning: Extended Abstract
- On The Transferability of Deep-Q Networks
- BOOK: Storing Algorithm-Invariant Episodes for Deep Reinforcement Learning
- Generating GPU Compiler Heuristics using Reinforcement Learning
- Accelerated Target Updates for Q-learning
- Reinforcement Explanation Learning
- SAVER: Safe Learning-Based Controller for Real-Time Voltage Regulation
- Building Intelligent Autonomous Navigation Agents
- A Threshold-based Scheme for Reinforcement Learning in Neural Networks
- Learning Efficient and Effective Exploration Policies with Counterfactual Meta Policy
- Unsupervised Program Synthesis for Images By Sampling Without Replacement
- Automatic, Dynamic, and Nearly Optimal Learning Rate Specification by Local Quadratic Approximation
- Uniform State Abstraction For Reinforcement Learning
- Searching with Opponent-Awareness
- Planning with Expectation Models for Control
- Multitasking Inhibits Semantic Drift
- Evolution of Q Values for Deep Q Learning in Stable Baselines
- Automatic low-bit hybrid quantization of neural networks through meta learning
- Differentiable Robust LQR Layers
- Two-stage training algorithm for AI robot soccer
- VisualEnv: visual Gym environments with Blender
- Deep Deterministic Path Following
- The Atari Data Scraper
- Eden: A Unified Environment Framework for Booming Reinforcement Learning Algorithms
- A Bayesian Approach to Reinforcement Learning of Vision-Based Vehicular Control
- When Autonomous Systems Meet Accuracy and Transferability through AI: A Survey
- Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text
- State Representation Learning from Demonstration
- State Constrained Stochastic Optimal Control Using LSTMs
- A Dynamics Perspective of Pursuit-Evasion Games of Intelligent Agents with the Ability to Learn
- Simultaneous Navigation and Construction Benchmarking Environments
- Robust Reinforcement Learning under model misspecification
- Thompson Sampling via Local Uncertainty
- Learning robust driving policies without online exploration
- Neural Networks and Denotation
- libGroomRL: Reinforcement Learning for Jets
- Unbiased Deep Reinforcement Learning: A General Training Framework for Existing and Future Algorithms
- AITuning: Machine Learning-based Tuning Tool for Run-Time Communication Libraries
- A Smart Sliding Chinese Pinyin Input Method Editor on Touchscreen
- Genome Variant Calling with a Deep Averaging Network
- Application of Deep Q-Network in Portfolio Management
- Combine PPO with NES to Improve Exploration
- DefogGAN: Predicting Hidden Information in the StarCraft Fog of War with Generative Adversarial Nets
- Communication Efficient Parallel Reinforcement Learning
- Domain Knowledge Integration By Gradient Matching For Sample-Efficient Reinforcement Learning
- Learning Memory-Dependent Continuous Control from Demonstrations
- Improving Automated Visual Fault Detection by Combining a Biologically Plausible Model of Visual Attention with Deep Learning
- Distantly Supervised Question Parsing
- Collaborative creativity with Monte-Carlo Tree Search and Convolutional Neural Networks
- Deep Reinforcement Learning for High Level Character Control
- A Novel Update Mechanism for Q-Networks Based On Extreme Learning Machines
- Deep Reinforcement Learning for Backup Strategies against Adversaries
- Research on Autonomous Maneuvering Decision of UCAV based on Approximate Dynamic Programming
- Randomized Policy Learning for Continuous State and Action MDPs
- From proprioception to long-horizon planning in novel environments: A hierarchical RL model
- Continuous Control for High-Dimensional State Spaces: An Interactive Learning Approach
- Learning to Grasp from 2.5D images: a Deep Reinforcement Learning Approach
- Fast Approximate Solutions using Reinforcement Learning for Dynamic Capacitated Vehicle Routing with Time Windows
- Reusability and Transferability of Macro Actions for Reinforcement Learning
- Logical Team Q-learning: An approach towards factored policies in cooperative MARL
- Safety-guaranteed Reinforcement Learning based on Multi-class Support Vector Machine
- Robby is Not a Robber (anymore): On the Use of Institutions for Learning Normative Behavior
- A Study of State Aliasing in Structured Prediction with RNNs
- Interactive Lungs Auscultation with Reinforcement Learning Agent
- Integration of Imitation Learning using GAIL and Reinforcement Learning using Task-achievement Rewards via Probabilistic Graphical Model
- A Methodology for the Development of RL-Based Adaptive Traffic Signal Controllers
- Policy Evaluation and Seeking for Multi-Agent Reinforcement Learning via Best Response
- Off-Policy Self-Critical Training for Transformer in Visual Paragraph Generation
- DCNNs: A Transfer Learning comparison of Full Weapon Family threat detection for Dual-Energy X-Ray Baggage Imagery
- Hippocampal representations emerge when training recurrent neural networks on a memory dependent maze navigation task
- Understanding in Artificial Intelligence
- Nonparametric Additive Value Functions: Interpretable Reinforcement Learning with an Application to Surgical Recovery
- Knowledge accumulating: The general pattern of learning
- Deep Reinforcement Learning Models Predict Visual Responses in the Brain: A Preliminary Result
- Auto-CASH: Autonomous Classification Algorithm Selection with Deep Q-Network
- Cognitive Radio Network Throughput Maximization with Deep Reinforcement Learning
- Solving Interactive Fiction Games via Partial Evaluation and Bounded Model Checking
- Playing Go without Game Tree Search Using Convolutional Neural Networks
- A Reinforcement Learning Approach to the View Planning Problem
- Training an Interactive Helper
- Lineage Evolution Reinforcement Learning
- Meta Reinforcement Learning with Distribution of Exploration Parameters Learned by Evolution Strategies
- Data-Efficient Methods for Dialogue Systems
- Challenging common bolus advisor for self-monitoring type-I diabetes patients using Reinforcement Learning
- Reinforcement Learning with Neural Networks for Quantum Multiple Hypothesis Testing
- Performance and Resilience of Cyber-Physical Control Systems with Reactive Attack Mitigation
- Fine-Grained AutoAugmentation for Multi-Label Classification
- Exploring Variational Deep Q Networks
- Deep Reinforcement Learning for Tactile Robotics: Learning to Type on a Braille Keyboard
- Architecting and Visualizing Deep Reinforcement Learning Models
- Understanding Information Processing in Human Brain by Interpreting Machine Learning Models
- Joint Perception and Control as Inference with an Object-based Implementation
- Language Inference with Multi-head Automata through Reinforcement Learning
- Optimising Stochastic Routing for Taxi Fleets with Model Enhanced Reinforcement Learning
- Playing Catan with Cross-dimensional Neural Network
- Efficient RDF Graph Storage based on Reinforcement Learning
- Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime
- Learning Agent for a Heat-Pump Thermostat With a Set-Back Strategy Using Model-Free Reinforcement Learning
- Intelligent Replication Management for HDFS Using Reinforcement Learning
- Reinforcement learning with distance-based incentive/penalty (DIP) updates for highly constrained industrial control systems
- Non-Asymptotic Analysis of Monte Carlo Tree Search
- Biomechanic Posture Stabilisation via Iterative Training of Multi-policy Deep Reinforcement Learning Agents
- Reinforcement Learning for Robust Missile Autopilot Design
- Inverse Policy Evaluation for Value-based Sequential Decision-making
- Scheduling and Power Control for Wireless Multicast Systems via Deep Reinforcement Learning
- Learning to Expand: Reinforced Pseudo-relevance Feedback Selection for Information-seeking Conversations
- On-Demand Video Dispatch Networks: A Scalable End-to-End Learning Approach
- Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
- A General Framework for Charger Scheduling Optimization Problems
- Discovering hierarchies using Imitation Learning from hierarchy aware policies
- Safety Aware Reinforcement Learning (SARL)
- Adaptive ROI Generation for Video Object Segmentation Using Reinforcement Learning
- RLCache: Automated Cache Management Using Reinforcement Learning
- How to Train your Quadrotor: A Framework for Consistently Smooth and Responsive Flight Control via Reinforcement Learning
- Experience enrichment based task independent reward model
- Reinforced Bit Allocation under Task-Driven Semantic Distortion Metrics
- Weighted Entropy Modification for Soft Actor-Critic
- Boosting Image Recognition with Non-differentiable Constraints
- Combining Off and On-Policy Training in Model-Based Reinforcement Learning
- Adaptive Neural Architectures for Recommender Systems
- Functional Regularization for Reinforcement Learning via Learned Fourier Features
- Importance of Environment Design in Reinforcement Learning: A Study of a Robotic Environment
- Investigation on the generalization of the Sampled Policy Gradient algorithm
- How to Organize your Deep Reinforcement Learning Agents: The Importance of Communication Topology
- Transfer Reinforcement Learning across Homotopy Classes
- Delayed Rewards Calibration via Reward Empirical Sufficiency
- Dex: Incremental Learning for Complex Environments in Deep Reinforcement Learning
- Deep Reinforcement Learning for Inquiry Dialog Policies with Logical Formula Embeddings
- CLUSE: Cross-Lingual Unsupervised Sense Embeddings
- Generalizing Decision Making for Automated Driving with an Invariant Environment Representation using Deep Reinforcement Learning
- Amortized Variational Deep Q Network
- MAPEL: Multi-Agent Pursuer-Evader Learning using Situation Report
- Reflecting After Learning for Understanding
- Decoupled Learning of Environment Characteristics for Safe Exploration
- Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement Learning
- Anomaly-resistant Graph Neural Networks via Neural Architecture Search
- DeepFoldit -- A Deep Reinforcement Learning Neural Network Folding Proteins
- Human-Level Control without Server-Grade Hardware
- SDN Flow Entry Management Using Reinforcement Learning
- RLINK: Deep Reinforcement Learning for User Identity Linkage
- Regulating Reward Training by Means of Certainty Prediction in a Neural Network-Implemented Pong Game
- Choosing the Best of Both Worlds: Diverse and Novel Recommendations through Multi-Objective Reinforcement Learning
- Cascaded LSTMs based Deep Reinforcement Learning for Goal-driven Dialogue
- Generative Actor-Critic: An Off-policy Algorithm Using the Push-forward Model
- Transferring Deep Reinforcement Learning with Adversarial Objective and Augmentation
- MSDF: A Deep Reinforcement Learning Framework for Service Function Chain Migration
- Playing optical tweezers with deep reinforcement learning: in virtual, physical and augmented environments
- Data-Efficient Reinforcement Learning for Malaria Control
- Make Bipedal Robots Learn How to Imitate
- From Persistent Homology to Reinforcement Learning with Applications for Retail Banking
- Multimedia Edge Computing
- End-to-End Race Driving with Deep Reinforcement Learning
- Easy Monotonic Policy Iteration
- To be a fast adaptive learner: using game history to defeat opponents
- Angrier Birds: Bayesian reinforcement learning
- Using reinforcement learning to design an AI assistantfor a satisfying co-op experience
- Room Clearance with Feudal Hierarchical Reinforcement Learning
- Utilizing Skipped Frames in Action Repeats via Pseudo-Actions
- Biological Blueprints for Next Generation AI Systems
- Implementation of Q Learning and Deep Q Network For Controlling a Self Balancing Robot Model
- Evening the Score: Targeting SARS-CoV-2 Protease Inhibition in Graph Generative Models for Therapeutic Candidates
- Scalable, Decentralized Multi-Agent Reinforcement Learning Methods Inspired by Stigmergy and Ant Colonies
- Dynamic Multichannel Access via Multi-agent Reinforcement Learning: Throughput and Fairness Guarantees
- Region Growing Curriculum Generation for Reinforcement Learning
- Solving Atari Games Using Fractals And Entropy
- Neural Optimization Kernel: Towards Robust Deep Learning
- Active Offline Policy Selection
- An Internal Covariate Shift Bounding Algorithm for Deep Neural Networks by Unitizing Layers' Outputs
- Towards Learning to Play Piano with Dexterous Hands and Touch
- Deep Reinforcement Learning for Complex Manipulation Tasks with Sparse Feedback
- MIME: Mutual Information Minimisation Exploration
- Learning Robust Controllers Via Probabilistic Model-Based Policy Search
- Identification and adaptive control of a high-contrast focal plane wavefront correction system
- Influence-Based Reinforcement Learning for Intrinsically-Motivated Agents
- Can Complex Collective Behaviour Be Generated Through Randomness, Memory and a Pinch of Luck?
- Implementing Inductive bias for different navigation tasks through diverse RNN attractors
- Koopman Spectrum Nonlinear Regulators and Efficient Online Learning
- A data-driven choice of misfit function for FWI using reinforcement learning
- Locality-Sensitive Experience Replay for Online Recommendation
- Memoryless Exact Solutions for Deterministic MDPs with Sparse Rewards
- Going Beyond Linear RL: Sample Efficient Neural Function Approximation
- Fast constraint satisfaction problem and learning-based algorithm for solving Minesweeper
- MobiFace: A Novel Dataset for Mobile Face Tracking in the Wild
- Deep RL Agent for a Real-Time Action Strategy Game
- REST: Performance Improvement of a Black Box Model via RL-based Spatial Transformation
- State Distribution-aware Sampling for Deep Q-learning
- MQGrad: Reinforcement Learning of Gradient Quantization in Parameter Server
- Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model
- Robust Dual View Deep Agent
- Weakly Supervised Video Summarization by Hierarchical Reinforcement Learning
- A Policy Efficient Reduction Approach to Convex Constrained Deep Reinforcement Learning
- Learning View and Target Invariant Visual Servoing for Navigation
- Theoretically Principled Deep RL Acceleration via Nearest Neighbor Function Approximation
- Adaptive perturbation adversarial training: based on reinforcement learning
- Bootstrapped Meta-Learning
- Optimal Network Control in Partially-Controllable Networks
- RL-NCS: Reinforcement learning based data-driven approach for nonuniform compressed sensing