8 papers
Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning
Omar Adalat, Edwin Hamel-De le Court, Francesco Belardinelli
Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action ma…
Robust Shielding for Safe Reinforcement Learning
Edwin Hamel-De le Court, Thom Badings, Alessandro Abate +2
Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs). However, existing shielding techniques typ…
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
Matteo Leonesi, Francesco Belardinelli, Flavio Corradini +1
Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences once monitoring is lifted. Curr…
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
Maksim Anisimov, Francesco Belardinelli, Matthew Wicker
Safety guarantees are a prerequisite to the deployment of reinforcement learning (RL) agents in safety-critical tasks. Often, deployment environments exhibit non-stationary dynamic…
Expressive Temporal Specifications for Reward Monitoring
Omar Adalat, Francesco Belardinelli
Specifying informative and dense reward functions remains a pivotal challenge in Reinforcement Learning, as it directly affects the efficiency of agent training. In this work, we h…
ATL*AS: An Automata-Theoretic Approach and Tool for the Verification of Strategic Abilities in Multi-Agent Systems
Sofia Garcia de Blas Garcia-Alcalde, Francesco Belardinelli
We present two novel symbolic algorithms for model checking the Alternating-time Temporal Logic ATL*, over both the infinite-trace and the finite-trace semantics. In particular, fo…