17 papers
Risk-Controlled Post-Processing of Decision Policies
Sunay Joshi, Tao Wang, Hamed Hassani +1
Predictive models are often deployed through existing decision policies that stakeholders are reluctant to change unless a risk constraint requires intervention. We study risk-cont…
Benchmarking Misuse Mitigation Against Covert Adversaries
Davis Brown, Mahdi Sabbaghi, Luze Sun +4
Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing safeguards by requesting help on small,…
Detecting Safety Violations Across Many Agent Traces
Adam Stein, Davis Brown, Hamed Hassani +2
To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex, and sometimes even adversar…
Safety Guardrails for LLM-Enabled Robots
Zachary Ravichandran, Alexander Robey, Vijay Kumar +2
Although the integration of large language models (LLMs) into robotics has unlocked transformative capabilities, it has also introduced significant safety concerns, ranging from av…
Contextual Safety Reasoning and Grounding for Open-World Robots
Zachary Ravichandran, David Snyder, Alexander Robey +3
Robots are increasingly operating in open-world environments where safe behavior depends on context: the same hallway may require different navigation strategies when crowded versu…
Watermarking Language Models with Error Correcting Codes
Patrick Chao, Yan Sun, Edgar Dobriban +1
Recent progress in large language models enables the creation of realistic machine-generated content. Watermarking is a promising approach to distinguish machine-generated text fro…