activity
20192024
most citedMaintenance Strategies for Sewer Pipes with Multi-State Degradation and Deep Reinforcement Learning

6 citations · 9 across the 2 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG20246 cited

Maintenance Strategies for Sewer Pipes with Multi-State Degradation and Deep Reinforcement Learning

Lisandro A. Jimenez-Roa, Thiago D. Simão, Zaharah Bukhsh +4

Large-scale infrastructure systems are crucial for societal welfare, and their effective management requires strategic forecasting and intervention methods that account for various…

cs.LG2023

Robust Active Measuring under Model Uncertainty

Merlijn Krale, Thiago D. Simão, Jana Tumova +1

Partial observability and uncertainty are common problems in sequential decision-making that particularly impede the use of formal models such as Markov decision processes (MDPs).…

cs.LG2023

Reinforcement Learning by Guided Safe Exploration

Qisong Yang, Thiago D. Simão, Nils Jansen +2

Safety is critical to broadening the application of reinforcement learning (RL). Often, we train RL agents in a controlled environment, such as a laboratory, before deploying them…

cs.LG2023

More for Less: Safe Policy Improvement With Stronger Performance Guarantees

Patrick Wienhöft, Marnix Suilen, Thiago D. Simão +3

In an offline reinforcement learning setting, the safe policy improvement (SPI) problem aims to improve the performance of a behavior policy according to which sample data has been…

cs.LG20223 cited

Safe Reinforcement Learning From Pixels Using a Stochastic Latent Representation

Yannick Hogewind, Thiago D. Simao, Tal Kachman +1

We address the problem of safe reinforcement learning from pixel observations. Inherent challenges in such settings are (1) a trade-off between reward optimization and adhering to…

cs.LG2019

Safe Policy Improvement with an Estimated Baseline Policy

Thiago D. Simão, Romain Laroche, Rémi Tachet des Combes

Previous work has shown the unreliability of existing algorithms in the batch Reinforcement Learning setting, and proposed the theoretically-grounded Safe Policy Improvement with B…