4 papers
Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance
Bohdan Turbal, Blossom Metevier, Max Springer +1
Adversarial attacks on large language models have limited practical impact despite extensive research. Optimization-based attacks such as Greedy Coordinate Gradient (GCG) (Zou et a…
Measuring Validity in LLM-based Resume Screening
Jane Castleman, Zeyu Shen, Blossom Metevier +2
Resume screening is perceived as a particularly suitable task for LLMs given their ability to analyze natural language; thus many entities rely on general purpose LLMs without furt…
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
Max Springer, Chung Peng Lee, Blossom Metevier +5
Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial…
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Yaswanth Chittepu, Blossom Metevier, Will Schwarzer +3
Existing approaches to language model alignment often treat safety as a tradeoff against helpfulness, which can lead to unacceptable responses in sensitive domains. To ensure relia…