8 papers
PolyAlign: Conditional Human-Distribution Alignment
L. D. M. S. Sai Teja, Ufaq Khan, Sathira Silva +2
Post-training methods such as supervised fine-tuning (SFT) and preference optimization typically align language models toward a single global assistant behavior. While effective fo…
Detecting Jailbreak Attempts in Clinical Training LLMs Through Automated Linguistic Feature Extraction
Tri Nguyen, Huy Hoang Bao Le, Lohith Srikanth Pentapalli +2
Detecting jailbreak attempts in clinical training large language models (LLMs) requires accurate modeling of linguistic deviations that signal unsafe or off-task user behavior. Pri…
VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
Trieu Hai Nguyen, Sivaswamy Akilesh
The rapid development research of Large Language Models (LLMs) based on transformer architectures raises key challenges, one of them being the task of distinguishing between human-…
Distribution Matching via Generalized Consistency Models
Sagar Shrestha, Rajesh Shrestha, Tri Nguyen +1
Recent advancement in generative models have demonstrated remarkable performance across various data modalities. Beyond their typical use in data synthesis, these models play a cru…
LLM-as-a-Fuzzy-Judge: Fine-Tuning Large Language Models as a Clinical Evaluation Judge with Fuzzy Logic
Weibing Zheng, Laurah Turner, Jess Kropczynski +3
Clinical communication skills are critical in medical education, and practicing and assessing clinical communication skills on a scale is challenging. Although LLM-powered clinical…
Jailbreak Detection in Clinical Training LLMs Using Feature-Based Predictive Models
Tri Nguyen, Lohith Srikanth Pentapalli, Magnus Sieverding +11
Jailbreaking in Large Language Models (LLMs) threatens their safe use in sensitive domains like education by allowing users to bypass ethical safeguards. This study focuses on dete…