10 papers
Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA
Samuel Tetteh, Udip Shrestha, Joshua R. Waite +1
Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, ins…
Seeing Before Colliding: Anticipatory Safe RL with Frozen Vision-Language Models
Samuel Tetteh, Cody Fleming
The cost signal that constrained-RL algorithms optimize against is almost always reactive: the simulator emits a non-zero cost only after a collision has begun, and the Lagrange mu…
COOPO: Cyclic Offline-Online Policy Optimization Algorithm
Qisai Liu, Zhanhong Jiang, Joshua Russell Waite +3
Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL demands prohibitive environment in…
Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance
Adam Haroon, Erick J. RodrÃguez-Seda, Cody Fleming +1
Safe reinforcement learning (RL) typically asks an agent should do. We ask it needs to act, and show that a single policy can jointly learn control…
LexiSafe: Offline Safe Reinforcement Learning with Lexicographic Safety-Reward Hierarchy
Hsin-Jung Yang, Zhanhong Jiang, Prajwal Koirala +3
Offline safe reinforcement learning (RL) is increasingly important for cyber-physical systems (CPS), where safety violations during training are unacceptable and only pre-collected…
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning
Prajwal Koirala, Cody Fleming
Generative models such as diffusion and flow-matching offer expressive policies for offline reinforcement learning (RL) by capturing rich, multimodal action distributions, but thei…