43 papers
Sharding Prevents LLM Oversight Failures and Adversarial Exploitation
Victor Akinwande, J. Zico Kolter, Aran Nayebi
Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidenc…
Compressed Sensing for Capability Localization in Large Language Models
Anna Bair, Yixuan Even Xu, Mingjie Sun +1
Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We show that Transformer architectur…
A New Framework for Cybersecurity Refusals in AI Agents
Eliot Krzysztof Jones, Mateusz Dziemian, Matt Fredrikson +1
Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domains like cybersecurity. Existin…
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
Jeremy Tien, Abishek Anand, Yu-Rou Tuan +3
As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding t…
Reevaluating Policy Gradient Methods for Imperfect-Information Games
Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour +6
In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed nu…
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
Jingchu Gai, Guanning Zeng, Christina Baek +4
Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving rea…