4 papers
Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes
Mohammed Alshaalan, Miguel R. D. Rodrigues
Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed perplexity-based detectors. We cast…
Robust LLM Performance Certification via Constrained Maximum Likelihood Estimation
Minghe Shen, Ananth Balashankar, Adam Fisch +2
The ability to rigorously estimate the failure rates of large language models (LLMs) is a prerequisite for their safe deployment. Currently, however, practitioners often face a tra…
Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models
Reem I. Masoud, Chen Feng, Shunta Asano +3
The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning datasets used for cultural adapt…
Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges
Chen Feng, Minghe Shen, Ananth Balashankar +2
Reliable certification of Large Language Models (LLMs)-verifying that failure rates are below a safety threshold-is critical yet challenging. While "LLM-as-a-Judge" offers scalabil…