Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Controlling Tool Use with Heading-Specific Activation Steering
Yuqi Chen, Vincent Siu, Yang Liu +2
Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily. We investigate whether too…
cs.AI2026
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
Vincent Siu, Nathan W. Henry, Nicholas Crispino +3
Current safety evaluations of language models rely on benchmark-based assessments that may miss localized vulnerabilities. We present RepIt, a simple and data-efficient framework f…
cs.AI2025
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
Vincent Siu, Nicholas Crispino, David Park +5
We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the genera…