activity
20242026
collaborators

9 papers

cs.LG2026

Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers

Farbod Faraji, Francesco Belardinelli

Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making. Yet high-fidelity simulations are prohibitively costly, and machi…

cs.AI2026

Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning

Tiberiu-Andrei Georgescu, Alexander W. Goodall, Dalal Alrajeh +2

Shielding is widely used to enforce safety in reinforcement learning (RL), ensuring that an agent's actions remain compliant with formal specifications. Classical shielding approac…

cs.LG2026

Safe Reinforcement Learning via Recovery-based Shielding with Gaussian Process Dynamics Models

Alexander W. Goodall, Francesco Belardinelli

Reinforcement learning (RL) is a powerful framework for optimal decision-making and control but often lacks provable guarantees for safety-critical applications. In this paper, we…

cs.MA2026

Convergence and Connectivity: Dynamics of Multi-Agent Q-Learning in Random Networks

Dan Leonte, Aamal Hussain, Raphael Huser +2

Beyond specific settings, many multi-agent learning algorithms fail to converge to an equilibrium solution, instead displaying complex, non-stationary behaviours such as recurrent…

cs.LG2026

Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning

Alexander W. Goodall, Edwin Hamel-De le Court, Francesco Belardinelli

Many reinforcement learning algorithms, particularly those that rely on return estimates for policy improvement, can suffer from poor sample efficiency and training instability due…

cs.LO2025

Synthesis of Safety Specifications for Probabilistic Systems

Gaspard Ohlmann, Edwin Hamel-De le Court, Francesco Belardinelli

Ensuring that agents satisfy safety specifications can be crucial in safety-critical environments. While methods exist for controller synthesis with safe temporal specifications, m…