9 papers
Capabilities Ain't All You Need: Measuring Propensities in AI
Daniel Romero-Alvarado, Fernando MartÃnez-Plumed, Lorenzo Pacchiardi +11
AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the te…
The science and practice of proportionality in AI risk evaluations
Carlos Mougan, Lauritz Morlock, Jair Aguirre +19
A global challenge in artificial intelligence (AI) regulation lies in achieving effective risk management without compromising innovation and technical progress. The European Union…
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
Miles Brundage, Noemi Dreksler, Aidan Homewood +45
We outline a vision for frontier AI auditing, which we define as rigorous third-party verification of frontier AI developers' safety and security claims, and evaluation of their sy…
Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
Yuchen Wu, Edward Sun, Kaijie Zhu +4
Large language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where us…
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
Chandler Smith, Marwa Abdulhai, Manfred Diaz +83
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with bo…
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
Stephen Fitz, Peter Romero, Steven Basart +2
Large Language Models increasingly mediate high-stakes interactions, intensifying research on their capabilities and safety. While recent work has shown that LLMs exhibit consisten…