6 citations · 6 across the 6 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning
Astrid Horn Brorholt, Maris F. L. Galesloot, Nils Jansen +2
Probabilistic shielding is a technique for safe reinforcement learning (RL). Typically, a static observer -- called the shield -- constrains the learning agent's actions to those f…
cs.LG2026
Perception-Based Beliefs for POMDPs with Visual Observations
Miriam Schäfers, Merlijn Krale, Thiago D. Simão +2
Partially observable Markov decision processes (POMDPs) are a principled planning model for sequential decision-making under uncertainty. Yet, real-world problems with high-dimensi…