Designing an efficient and equitable humanitarian supply chain dynamically via reinforcement learning
arXiv:2505.17439
Abstract
This study designs an efficient and equitable humanitarian supply chain dynamically by using reinforcement learning, PPO, and compared with heuristic algorithms. This study demonstrates the model of PPO always treats average satisfaction rate as the priority.