activity
20242026
most citedNavigating Demand Uncertainty in Container Shipping: Deep Reinforcement Learning for Enabling Adaptive and Feasible Master Stowage Planning

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

Weikai Wang, Erick Delage

Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restr…

cs.LG20261 cited

Navigating Demand Uncertainty in Container Shipping: Deep Reinforcement Learning for Enabling Adaptive and Feasible Master Stowage Planning

Jaike van Twiller, Yossiri Adulyasak, Erick Delage +2

Reinforcement learning (RL) has successfully solved various deterministic and stochastic planning problems. However, conventional RL struggles with complex real-world constraints,…

cs.LG2026

Risk-Aware Decision Making in Restless Bandits: Theory and Algorithms for Planning and Learning

Nima Akbarzadeh, Yossiri Adulyasak, Erick Delage

In restless bandits, a central agent is tasked with optimally distributing limited resources across several bandits (arms), with each arm being a Markov decision process. In this w…

cs.LG2025

Planning and Learning in Average Risk-aware MDPs

Weikai Wang, Erick Delage

For continuing tasks, average cost Markov decision processes have well-documented value and can be solved using efficient algorithms. However, it explicitly assumes that the agent…

cs.LG2025

Fair Resource Allocation in Weakly Coupled Markov Decision Processes

Xiaohui Tu, Yossiri Adulyasak, Nima Akbarzadeh +1

We consider fair resource allocation in sequential decision-making environments modeled as weakly coupled Markov decision processes, where resource constraints couple the action sp…

cs.LG2024

Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis

Jia Lin Hau, Erick Delage, Esther Derman +2

In Markov decision processes (MDPs), quantile risk measures such as Value-at-Risk are a standard metric for modeling RL agents' preferences for certain outcomes. This paper propose…