5 papers · 1 filter
Controlling Tool Use with Heading-Specific Activation Steering
Yuqi Chen, Vincent Siu, Yang Liu +2
Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily. We investigate whether too…
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
Vincent Siu, Manasi Sharma, Dawn Song +3
Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap wit…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
Vincent Siu, Nathan W. Henry, Nicholas Crispino +3
Current safety evaluations of language models rely on benchmark-based assessments that may miss localized vulnerabilities. We present RepIt, a simple and data-efficient framework f…
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
Vincent Siu, Nicholas Crispino, David Park +5
We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the genera…