7 papers
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
Kaiwen Zhou, Ahmed Elgohary, A S M Iftekhar +1
The ability of LLM agents to plan and invoke tools exposes them to new safety risks, making a comprehensive red-teaming system crucial for discovering vulnerabilities and ensuring…
A Definition of AGI
Dan Hendrycks, Dawn Song, Christian Szegedy +30
The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quant…
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
Satyapriya Krishna, Andy Zou, Rahul Gupta +6
The safety and alignment of Large Language Models (LLMs) are critical for their responsible deployment. Current evaluation methods predominantly focus on identifying and preventing…
Evaluating Language Model Reasoning about Confidential Information
Dylan Sam, Alexander Robey, Andy Zou +2
As language models are increasingly deployed as autonomous agents in high-stakes settings, ensuring that they reliably follow user-defined rules has become a critical safety concer…
Adversarial Attacks on Robotic Vision Language Action Models
Eliot Krzysztof Jones, Alexander Robey, Andy Zou +5
The emergence of vision-language-action models (VLAs) for end-to-end control is reshaping the field of robotics by enabling the fusion of multimodal sensory inputs at the billion-p…
Transferable Adversarial Attacks on Black-Box Vision-Language Models
Kai Hu, Weichen Yu, Li Zhang +5
Vision Large Language Models (VLLMs) are increasingly deployed to offer advanced capabilities on inputs comprising both text and images. While prior research has shown that adversa…