Medical Dead-ends and Learning to Identify High-risk States and Treatments
arXiv:2110.04186
Abstract
Machine learning has successfully framed many sequential decision making problems as either supervised prediction, or optimal decision-making policy identification via reinforcement learning. In data-constrained offline settings, both approaches may fail as they assume fully optimal behavior or rely on exploring alternatives that may not exist. We introduce an inherently different approach that identifies possible "dead-ends" of a state space. We focus on the condition of patients in the intensive care unit, where a "medical dead-end" indicates that a patient will expire, regardless of all potential future treatment sequences. We postulate "treatment security" as avoiding treatments with probability proportional to their chance of leading to dead-ends, present a formal proof, and frame discovery as an RL problem. We then train three independent deep neural models for automated state construction, dead-end discovery and confirmation. Our empirical results discover that dead-ends exist in real clinical data among septic patients, and further reveal gaps between secure treatments and those that were administered.
References in corpus (11)
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Behavior Regularized Offline Reinforcement Learning
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Continuous State-Space Models for Optimal Sepsis Treatment - a Deep Reinforcement Learning Approach
- Critic Regularized Regression
- An Improved Multi-Output Gaussian Process RNN with Real-Time Validation for Early Sepsis Detection
- Improving Sepsis Treatment Strategies by Combining Deep and Kernel-Based Reinforcement Learning
- Clinician-in-the-Loop Decision Making: Reinforcement Learning with Near-Optimal Set-Valued Policies
- Optimizing Sequential Medical Treatments with Auto-Encoding Heuristic Search in POMDPs
- S4RL: Surprisingly Simple Self-Supervision for Offline Reinforcement Learning