Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
A Systematic Investigation of RL-Jailbreaking in LLMs
Montaser Mohammedalamen, Kevin Roice, Reginald McLean +1
The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversarial jailbreaking, the strateg…
cs.LG2025
Multi-Task Reinforcement Learning Enables Parameter Scaling
Reginald McLean, Evangelos Chatzaroulas, Jordan Terry +3
Multi-task reinforcement learning (MTRL) aims to endow a single agent with the ability to perform well on multiple tasks. Recent works have focused on developing novel sophisticate…