1 paper · 1 filter
Tri Nguyen, Huy Hoang Bao Le, Lohith Srikanth Pentapalli +2
Detecting jailbreak attempts in clinical training large language models (LLMs) requires accurate modeling of linguistic deviations that signal unsafe or off-task user behavior. Pri…