2 papers
cs.LG2026
Position: Deployed Reinforcement Learning should be Continual
Parnian Behdin, Kevin Roice, Golnaz Mesbahi
Reinforcement Learning (RL) has received increasing attention and adoption in real-world use cases. Most of these systems follow a train-then-fix paradigm, where trained agents do…
cs.LG2026
A Systematic Investigation of RL-Jailbreaking in LLMs
Montaser Mohammedalamen, Kevin Roice, Reginald McLean +1
The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversarial jailbreaking, the strateg…