Discovering an Aid Policy to Minimize Student Evasion Using Offline Reinforcement Learning
arXiv:2104.10258 · doi:10.1109/IJCNN52387.2021.9534159
Abstract
High dropout rates in tertiary education expose a lack of efficiency that causes frustration of expectations and financial waste. Predicting students at risk is not enough to avoid student dropout. Usually, an appropriate aid action must be discovered and applied in the proper time for each student. To tackle this sequential decision-making problem, we propose a decision support method to the selection of aid actions for students using offline reinforcement learning to support decision-makers effectively avoid student dropout. Additionally, a discretization of student's state space applying two different clustering methods is evaluated. Our experiments using logged data of real students shows, through off-policy evaluation, that the method should achieve roughly 1.0 to 1.5 times as much cumulative reward as the logged policy. So, it is feasible to help decision-makers apply appropriate aid actions and, possibly, reduce student dropout.
8 pages, 6 figures, accepted for publication in 2021 International Joint Conference on Neural Networks (IJCNN 2021)
References in corpus (11)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
- Doubly Robust Policy Evaluation and Learning
- Behavior Regularized Offline Reinforcement Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
- Deep Reinforcement Learning for Sepsis Treatment
- Detecting Troll Behavior via Inverse Reinforcement Learning: A Case Study of Russian Trolls in the 2016 US Election