2 papers
cs.CL2025
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights
Alexandra Abbas, Celia Waggoner, Justin Olive
AI evaluations have become critical tools for assessing large language model capabilities and safety. This paper presents practical insights from eight months of maintaining $inspe…
cs.CL2025
Latent Adversarial Training Improves the Representation of Refusal
Alexandra Abbas, Nora Petrova, Helios Ael Lyons +1
Recent work has shown that language models' refusal behavior is primarily encoded in a single direction in their latent space, making it vulnerable to targeted attacks. Although La…