collaborators

10 papers

cs.LG2026

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

Samuel Tetteh, Udip Shrestha, Joshua R. Waite +1

Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, ins…

cs.LG2026

Seeing Before Colliding: Anticipatory Safe RL with Frozen Vision-Language Models

Samuel Tetteh, Cody Fleming

The cost signal that constrained-RL algorithms optimize against is almost always reactive: the simulator emits a non-zero cost only after a collision has begun, and the Lagrange mu…

cs.LG2026

COOPO: Cyclic Offline-Online Policy Optimization Algorithm

Qisai Liu, Zhanhong Jiang, Joshua Russell Waite +3

Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL demands prohibitive environment in…

cs.LG2026

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

Adam Haroon, Erick J. Rodríguez-Seda, Cody Fleming +1

Safe reinforcement learning (RL) typically asks an agent should do. We ask it needs to act, and show that a single policy can jointly learn control…

cs.LG2026

LexiSafe: Offline Safe Reinforcement Learning with Lexicographic Safety-Reward Hierarchy

Hsin-Jung Yang, Zhanhong Jiang, Prajwal Koirala +3

Offline safe reinforcement learning (RL) is increasingly important for cyber-physical systems (CPS), where safety violations during training are unacceptable and only pre-collected…

cs.LG2026

Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning

Prajwal Koirala, Cody Fleming

Generative models such as diffusion and flow-matching offer expressive policies for offline reinforcement learning (RL) by capturing rich, multimodal action distributions, but thei…