3 papers
cs.AI2026
Detecting Jailbreak Attempts in Clinical Training LLMs Through Automated Linguistic Feature Extraction
Tri Nguyen, Huy Hoang Bao Le, Lohith Srikanth Pentapalli +2
Detecting jailbreak attempts in clinical training large language models (LLMs) requires accurate modeling of linguistic deviations that signal unsafe or off-task user behavior. Pri…
cs.CR2025
A Gradient-Optimized TSK Fuzzy Framework for Explainable Phishing Detection
Lohith Srikanth Pentapalli, Jon Salisbury, Josette Riep +1
Phishing attacks represent an increasingly sophisticated and pervasive threat to individuals and organizations, causing significant financial losses, identity theft, and severe dam…
cs.CL2025
Jailbreak Detection in Clinical Training LLMs Using Feature-Based Predictive Models
Tri Nguyen, Lohith Srikanth Pentapalli, Magnus Sieverding +11
Jailbreaking in Large Language Models (LLMs) threatens their safe use in sensitive domains like education by allowing users to bypass ethical safeguards. This study focuses on dete…