2 papers
cs.AI2026
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
Vincent Siu, Nathan W. Henry, Nicholas Crispino +3
Current safety evaluations of language models rely on benchmark-based assessments that may miss localized vulnerabilities. We present RepIt, a simple and data-efficient framework f…
cs.AI2025
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
Vincent Siu, Nicholas Crispino, David Park +5
We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the genera…