4 papers
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
Chandler Smith, Marwa Abdulhai, Manfred Diaz +83
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with bo…
Governing Automated Strategic Intelligence
Nicholas Kruus, Madhavendra Thakur, Adam Khoja +24
Military and economic strategic competitiveness between nation-states will increasingly be defined by the capability and cost of their frontier artificial intelligence models. Amon…
Probing and Steering Evaluation Awareness of Language Models
Jord Nguyen, Khiem Hoang, Carlo Leonardo Attubato +1
Language models can distinguish between testing and deployment phases -- a capability known as evaluation awareness. This has significant safety and policy implications, potentiall…
DarkBench: Benchmarking Dark Patterns in Large Language Models
Esben Kran, Hieu Minh "Jord" Nguyen, Akash Kundu +3
We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language mode…