5 papers
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
Buyun Liang, Jinqi Luo, Liangzu Peng +6
Large language models (LLMs) achieve strong performance across many tasks but remain vulnerable to hallucinations, making it important to systematically evaluate their reliability…
Stateful Online Monitoring Catches Distributed Agent Attacks
Davis Brown, Samarth Bhargav, Arav Santhanam +7
Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection, attackers frequently distribu…
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
Mahdi Sabbaghi, George Pappas, Adel Javanmard +1
Supervised fine-tuning (SFT) provides the standard approach for teaching LLMs new behaviors from offline expert demonstrations. However, standard SFT uniformly fits all samples --…
MultiRisk: Multiple Risk Control via Iterative Score Thresholding
Sunay Joshi, Yan Sun, Hamed Hassani +1
As generative AI systems are increasingly deployed in real-world applications, regulating multiple dimensions of model behavior has become essential. We focus on test-time filterin…
Watermark Smoothing Attacks against Language Models
Hongyan Chang, Hamed Hassani, Reza Shokri
Watermarking is a key technique for detecting AI-generated text. In this work, we study its vulnerabilities and introduce the Smoothing Attack, a novel watermark removal method. By…