Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Tamper-Resistant Safeguards for Open-Weight LLMs
Rishub Tamirisa, Bhrugu Bharathi, Long Phan +12
Rapid advances in the capabilities of large language models (LLMs) have raised widespread concerns regarding their potential for malicious use. Open-weight LLMs present unique chal…
cs.LG2024
Improving Alignment and Robustness with Circuit Breakers
Andy Zou, Long Phan, Justin Wang +7
AI systems can take harmful actions and are highly vulnerable to adversarial attacks. We present an approach, inspired by recent advances in representation engineering, that interr…