2 papers
cs.CR2026
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
Saad Hossain, Tom Tseng, Punya Syon Pandey +8
As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, be…
cs.LG2025
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
Saad Hossain, Samanvay Vajpayee, Sirisha Rambhatla
As large language models (LLMs) become ubiquitous, parameter-efficient fine-tuning methods and safety-first defenses have proliferated rapidly. However, the number of approaches an…