4 papers
A Systematic Investigation of RL-Jailbreaking in LLMs
Montaser Mohammedalamen, Kevin Roice, Reginald McLean +1
The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversarial jailbreaking, the strateg…
Learning to Be Cautious
Montaser Mohammedalamen, Dustin Morrill, Alexander Sieusahai +2
A key challenge in the field of reinforcement learning is to develop agents that behave cautiously in novel situations. It is generally impossible to anticipate all situations that…
Generalization in Monitored Markov Decision Processes (Mon-MDPs)
Montaser Mohammedalamen, Michael Bowling
Reinforcement learning (RL) typically models the interaction between the agent and environment as a Markov decision process (MDP), where the rewards that guide the agent's behavior…
Monitored Markov Decision Processes
Simone Parisi, Montaser Mohammedalamen, Alireza Kazemipour +2
In reinforcement learning (RL), an agent learns to perform a task by interacting with an environment and receiving feedback (a numerical reward) for its actions. However, the assum…