papers

Publications (57)

eess.SY2021

Bias Correction in Deterministic Policy Gradient Using Robust MPC

Arash Bahari Kordabad, Hossein Nejatbakhsh Esfahani, Sebastien Gros

In this paper, we discuss the deterministic policy gradient using the Actor-Critic methods based on the linear compatible advantage function approximator, where the input spaces ar…

math.OC2025

Differentiable Nonlinear Model Predictive Control

Jonathan Frey, Katrin Baumgärtner, Gianluca Frison +5

The efficient computation of parametric solution sensitivities is a key challenge in the integration of learning-enhanced methods with nonlinear model predictive control (MPC), as…

eess.SY2023

Interpretable Battery Cycle Life Range Prediction Using Early Degradation Data at Cell Level

Huang Zhang, Yang Su, Faisal Altaf +2

Battery cycle life prediction using early degradation data has many potential applications throughout the battery product life cycle. For that reason, various data-driven methods h…

cs.CE2025

Efficient Multi-Objective Constrained Bayesian Optimization of Bridge Girder

Heine Havneraas Røstum, Joseph Morlier, Sebastien Gros +1

The buildings and construction sector is a significant source of greenhouse gas emissions, with cement production alone contributing 7~\% of global emissions and the industry as a…

math.OC2022

Recursive Feasibility of Stochastic Model Predictive Control with Mission-Wide Probabilistic Constraints

Kai Wang, Sebastien Gros

This paper is concerned with solving chance-constrained finite-horizon optimal control problems, with a particular focus on the recursive feasibility issue of stochastic model pred…

eess.SY2021

Reinforcement Learning based on MPC/MHE for Unmodeled and Partially Observable Dynamics

Hossein Nejatbakhsh Esfahani, Arash Bahari Kordabad, Sebastien Gros

This paper proposes an observer-based framework for solving Partially Observable Markov Decision Processes (POMDPs) when an accurate model is not available. We first propose to use…

eess.SY2020

Safe Reinforcement Learning via Projection on a Safe Set: How to Achieve Optimality?

Sebastien Gros, Mario Zanon, Alberto Bemporad

For all its successes, Reinforcement Learning (RL) still struggles to deliver formal guarantees on the closed-loop behavior of the learned policy. Among other things, guaranteeing…

eess.SY2021

Optimization of the Model Predictive Control Meta-Parameters Through Reinforcement Learning

Eivind Bøhn, Sebastien Gros, Signe Moe +1

Model predictive control (MPC) is increasingly being considered for control of fast systems and embedded applications. However, the MPC has some significant challenges for such sys…

eess.SY2026

Solving Markov Decision Processes with Future Information via MPC

Shambhuraj Sawant, Akhil S Anand, Dirk Reinhardt +1

Model Predictive Control (MPC) is widely used in industrial and robotic systems for enforcing constraints and embedding domain knowledge through finite-horizon optimization-based p…

eess.SY2019

Towards Safe Reinforcement Learning Using NMPC and Policy Gradients: Part II - Deterministic Case

Sebastien Gros, Mario Zanon

In this paper, we present a methodology to deploy the deterministic policy gradient method, using actor-critic techniques, when the optimal policy is approximated using a parametri…

math.OC2024

Optimality Conditions for Model Predictive Control: Rethinking Predictive Model Design

Akhil S Anand, Arash Bahari Kordabad, Mario Zanon +1

Optimality is a critical aspect of Model Predictive Control (MPC), especially in economic MPC. However, achieving optimality in MPC presents significant challenges, and may even be…

math.OC2017

A Feasibility-Enforcing Primal-Decomposition SQP Algorithm for Optimal Vehicle Coordination

Mario Zanon, Robert Hult, Sebastien Gros +1

In this paper we consider the problem of coordinating autonomous vehicles approaching an intersection. We cast the problem in the distributed optimisation framework and propose an…

eess.SY2022

Policy Gradient Reinforcement Learning for Uncertain Polytopic LPV Systems based on MHE-MPC

Hossein Nejatbakhsh Esfahani, Sebastien Gros

In this paper, we propose a learning-based Model Predictive Control (MPC) approach for the polytopic Linear Parameter-Varying (LPV) systems with inexact scheduling parameters (as e…

cs.LG2021

MPC-based Reinforcement Learning for Economic Problems with Application to Battery Storage

Arash Bahari Kordabad, Wenqi Cai, Sebastien Gros

In this paper, we are interested in optimal control problems with purely economic costs, which often yield optimal policies having a (nearly) bang-bang structure. We focus on polic…

eess.SY2024

Data-Driven Domestic Flexible Demand: Observations from experiments in cold climate

Dirk Reinhardt, Wenqi Cai, Sebastien Gros

In this chapter, we report on our experience with domestic flexible electric energy demand based on a regular commercial (HVAC)-based heating system in a house. Our focus is on inv…

cs.LG2025

Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization Strategies

Runze Yan, Xun Shen, Akifumi Wachi +3

When applying offline reinforcement learning (RL) in healthcare scenarios, the out-of-distribution (OOD) issues pose significant risks, as inappropriate generalization beyond clini…

eess.SY2024

Economic Model Predictive Control as a Solution to Markov Decision Processes

Dirk Reinhardt, Akhil S. Anand, Shambhuraj Sawant +1

Markov Decision Processes (MDPs) offer a fairly generic and powerful framework to discuss the notion of optimal policies for dynamic systems, in particular when the dynamics are st…

eess.SY2022

Bridging the gap between QP-based and MPC-based RL

Shambhuraj Sawant, Sebastien Gros

Reinforcement learning methods typically use Deep Neural Networks to approximate the value functions and policies underlying a Markov Decision Process. Unfortunately, DNN-based RL…

eess.SY2021

Approximate Robust NMPC using Reinforcement Learning

Hossein Nejatbakhsh Esfahani, Arash Bahari Kordabad, Sebastien Gros

We present a Reinforcement Learning-based Robust Nonlinear Model Predictive Control (RL-RNMPC) framework for controlling nonlinear systems in the presence of disturbances and uncer…

eess.SY2020

Reinforcement Learning for Mixed-Integer Problems Based on MPC

Sebastien Gros, Mario Zanon

Model Predictive Control has been recently proposed as policy approximation for Reinforcement Learning, offering a path towards safe and explainable Reinforcement Learning. This ap…

eess.SY2021

Verification of Dissipativity and Evaluation of Storage Function in Economic Nonlinear MPC using Q-Learning

Arash Bahari Kordabad, Sebastien Gros

In the Economic Nonlinear Model Predictive (ENMPC) context, closed-loop stability relates to the existence of a storage function satisfying a dissipation inequality. Finding the st…

eess.SY2021

A Semi-Distributed Interior Point Algorithm for Optimal Coordination of Automated Vehicles at Intersections

Robert Hult, Mario Zanon, Sebastien Gros +1

In this paper, we consider the optimal coordination of automated vehicles at intersections under fixed crossing orders. We formulate the problem using direct optimal control and ex…

cs.LG2025

Quasi-Newton Compatible Actor-Critic for Deterministic Policies

Arash Bahari Kordabad, Dean Brandner, Sebastien Gros +2

In this paper, we propose a second-order deterministic actor-critic framework in reinforcement learning that extends the classical deterministic policy gradient method to exploit c…

eess.SY2026

Sample-Efficient Model-Free Policy Gradient Methods for Stochastic LQR via Robust Linear Regression

Bowen Song, Sebastien Gros, Andrea Iannelli

Policy gradient algorithms are widely used in reinforcement learning and belong to the class of approximate dynamic programming methods. This paper studies two key policy gradient…

eess.SY2026

Computationally efficient Gauss-Newton reinforcement learning for model predictive control

Dean Brandner, Sebastien Gros, Sergio Lucia

Model predictive control (MPC) is widely used in process control due to its interpretability and ability to handle constraints. As a parametric policy in reinforcement learning (RL…

cs.AI2025

All AI Models are Wrong, but Some are Optimal

Akhil S Anand, Shambhuraj Sawant, Dirk Reinhardt +1

AI models that predict the future behavior of a system (a.k.a. predictive AI models) are central to intelligent decision-making. However, decision-making using predictive AI models…

cs.AI2025

CORL: Reinforcement Learning of MILP Policies Solved via Branch and Bound

Akhil S Anand, Elias Aarekol, Martin Mziray Dalseg +2

Combinatorial sequential decision making problems are typically modeled as mixed integer linear programs (MILPs) and solved via branch and bound (B&B) algorithms. The inherent diff…

eess.SY2023

Equivalence of Optimality Criteria for Markov Decision Process and Model Predictive Control

Arash Bahari Kordabad, Mario Zanon, Sebastien Gros

This paper shows that the optimal policy and value functions of a Markov Decision Process (MDP), either discounted or not, can be captured by a finite-horizon undiscounted Optimal…

cs.LG2026

Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?

Hannah Markgraf, Shambhuraj Sawant, Hanna Krasowski +3

Projection-based safety filters, which modify unsafe actions by mapping them to the closest safe alternative, are widely used to enforce safety constraints in reinforcement learnin…

eess.SY2024

Personalized Dynamic Pricing Policy for Electric Vehicles: Reinforcement learning approach

Sangjun Bae, Balazs Kulcsar, Sebastien Gros

With the increasing number of fast-electric vehicle charging stations (fast-EVCSs) and the popularization of information technology, electricity price competition between fast-EVCS…

eess.SY2020

Optimal model-based trajectory planning with static polygonal constraints

Andreas B. Martinsen, Anastasios M. Lekkas, Sebastien Gros

The main contribution of this paper is a novel method for planning globally optimal trajectories for dynamical systems subject to polygonal constraints. The proposed method is a hy…

cs.LG2022

Quasi-Newton Iteration in Deterministic Policy Gradient

Arash Bahari Kordabad, Hossein Nejatbakhsh Esfahani, Wenqi Cai +1

This paper presents a model-free approximation for the Hessian of the performance of deterministic policies to use in the context of Reinforcement Learning based on Quasi-Newton st…

eess.SY2020

Optimization of the Model Predictive Control Update Interval Using Reinforcement Learning

Eivind Bøhn, Sebastien Gros, Signe Moe +1

In control applications there is often a compromise that needs to be made with regards to the complexity and performance of the controller and the computational resources that are…

eess.SY2025

Data-Driven Distributionally Robust Control for Interacting Agents under Logical Constraints

Arash Bahari Kordabad, Eleftherios E. Vlahakis, Lars Lindemann +3

In this paper, we propose a distributionally robust control synthesis for an agent with stochastic dynamics that interacts with other agents under uncertainties and constraints exp…

math.OC2025

Optimization-based Coordination of Traffic Lights and Automated Vehicles at Intersections

Azita Dabiri, Giray Önür, Sebastien Gros +1

This paper tackles the challenge of coordinating traffic lights and automated vehicles at signalized intersections, formulated as a constrained finite-horizon optimal control probl…

eess.SY2025

MPC4RL -- A Software Package for Reinforcement Learning based on Model Predictive Control

Dirk Reinhardt, Katrin Baumgärnter, Jonathan Frey +2

In this paper, we present an early software integrating Reinforcement Learning (RL) with Model Predictive Control (MPC). Our aim is to make recent theoretical contributions from th…

eess.SY2021

Reinforcement Learning of the Prediction Horizon in Model Predictive Control

Eivind Bøhn, Sebastien Gros, Signe Moe +1

Model predictive control (MPC) is a powerful trajectory optimization control technique capable of controlling complex nonlinear systems while respecting system constraints and ensu…

eess.SY2023

Once upon a time step: A closed-loop approach to robust MPC design

Anilkumar Parsi, Marcell Bartos, Amber Srivastava +2

A novel perspective on the design of robust model predictive control (MPC) methods is presented, whereby closed-loop constraint satisfaction is ensured using recursive feasibility…

cs.LG2026

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

Yuepeng Wang, Ken Kawano, Yongqi Zhou +11

Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performe…

eess.SY2019

Towards Safe Reinforcement Learning Using NMPC and Policy Gradients: Part I - Stochastic case

Sebastien Gros, Mario Zanon

We present a methodology to deploy the stochastic policy gradient method, using actor-critic techniques, when the optimal policy is approximated using a parametric optimization pro…

cs.LG2026

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

Xun Shen, Yuepeng Wang, Akifumi Wachi +13

Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical inte…

cs.AI2023

Deep active learning for nonlinear system identification

Erlend Torje Berg Lundby, Adil Rasheed, Ivar Johan Halvorsen +3

The exploding research interest for neural networks in modeling nonlinear dynamical systems is largely explained by the networks' capacity to model complex input-output relations d…

eess.SY2021

Backstepping-based Integral Sliding Mode Control with Time Delay Estimation for Autonomous Underwater Vehicles

Hossein Nejatbakhsh Esfahani, Behdad Aminian, Esten Ingar Grøtli +1

The aim of this paper is to propose a high performance control approach for trajectory tracking of Autonomous Underwater Vehicles (AUVs). However, the controller performance can be…

cs.LG2024

Flipping-based Policy for Chance-Constrained Markov Decision Processes

Xun Shen, Shuo Jiang, Akifumi Wachi +2

Safe reinforcement learning (RL) is a promising approach for many real-world decision-making problems where ensuring safety is a critical necessity. In safe RL research, while expe…

eess.SY2024

Application of Soft Actor-Critic Algorithms in Optimizing Wastewater Treatment with Time Delays Integration

Esmaeel Mohammadi, Daniel Ortiz-Arroyo, Aviaja Anna Hansen +4

Wastewater treatment plants face unique challenges for process control due to their complex dynamics, slow time constants, and stochastic delays in observations and actions. These…

eess.SY2022

Functional Stability of Discounted Markov Decision Processes Using Economic MPC Dissipativity Theory

Arash Bahari Kordabad, Sebastien Gros

This paper discusses the functional stability of closed-loop Markov Chains under optimal policies resulting from a discounted optimality criterion, forming Markov Decision Processe…

math.OC2024

Data-Driven Predictive Control and MPC: Do we achieve optimality?

Akhil S Anand, Shambhuraj Sawant, Dirk Reinhardt +1

In this paper, we explore the interplay between Predictive Control and closed-loop optimality, spanning from Model Predictive Control to Data-Driven Predictive Control. Predictive…

eess.SY2021

MPC-based Reinforcement Learning for a Simplified Freight Mission of Autonomous Surface Vehicles

Wenqi Cai, Arash B. Kordabad, Hossein N. Esfahani +2

In this work, we propose a Model Predictive Control (MPC)-based Reinforcement Learning (RL) method for Autonomous Surface Vehicles (ASVs). The objective is to find an optimal polic…

cs.LG2025

Bi-Level Reinforcement Learning Pathway for Sim-to-Real Optimality

Akhil S Anand, Shambhuraj Sawant, Paavo Parmas +3

Training Reinforcement Learning (RL) policies using simulation models before deployment in real-world environments is a common strategy when real-world interaction is expensive. Th…

eess.SY2020

Combining system identification with reinforcement learning-based MPC

Andreas B. Martinsen, Anastasios M. Lekkas, Sebastien Gros

In this paper we propose and compare methods for combining system identification (SYSID) and reinforcement learning (RL) in the context of data-driven model predictive control (MPC…

eess.SY2019

Autonomous docking using direct optimal control

Andreas B. Martinsen, Anastasios M. Lekkas, Sebastien Gros

We propose a method for performing autonomous docking of marine vessels using numerical optimal control. The task is framed as a dynamic positioning problem, with the addition of s…

eess.SY2023

Learning-based MPC from Big Data Using Reinforcement Learning

Shambhuraj Sawant, Akhil S Anand, Dirk Reinhardt +1

This paper presents an approach for learning Model Predictive Control (MPC) schemes directly from data using Reinforcement Learning (RL) methods. The state-of-the-art learning meth…

math.OC2026

Mission-Aligned Learning-Informed Control of Autonomous Systems: Formulation and Foundations

Vyacheslav Kungurtsev, Monicah Cherop Naibei, Gustav Sir +4

Research, innovation and practical capital investment have been increasing rapidly toward the realization of autonomous physical agents. This includes industrial and service robots…

eess.SY2025

Probabilistic reachable sets of stochastic nonlinear systems with contextual uncertainties

Xun Shen, Ye Wang, Kazumune Hashimoto +2

Validating and controlling safety-critical systems in uncertain environments necessitates probabilistic reachable sets of future state evolutions. The existing methods of computing…

eess.SY2025

Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification

Rudolf Reiter, Jasper Hoffmann, Dirk Reinhardt +6

The fields of MPC and RL consider two successful control techniques for Markov decision processes. Both approaches are derived from similar fundamental principles, and both are wid…

eess.SY2026

Uncertainty Propagation under Residual Disturbances: A Smart-Home Case Study

Guanru Pan, Dirk Reinhardt, Sebastien Gros +1

This paper presents a data-driven framework for uncertainty propagation under unmeasured or statistically unmodeled (unstructured) disturbances. We consider residual disturbances,…

math.OC2022

Solving Mission-Wide Chance-Constrained Optimal Control Using Dynamic Programming

Kai Wang, Sebastien Gros

This paper aims to provide a Dynamic Programming (DP) approach to solve the Mission-Wide Chance-Constrained Optimal Control Problems (MWCC-OCP). The mission-wide chance constraint…