3 citations · 3 across the 3 of their papers we have counts for
6 papers
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…
Documenting Deployment with Fabric: A Repository of Real-World AI Governance
Mackenzie Jorgensen, Kendall Brogle, Katherine M. Collins +10
Artificial intelligence (AI) is increasingly integrated into society, from financial services and traffic management to creative writing. Academic literature on the deployment of A…
Training language models to be warm and empathetic makes them less reliable and more sycophantic
Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher
Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and compani…
Promising Topics for U.S.-China Dialogues on AI Risks and Governance
Saad Siddiqui, Lujain Ibrahim, Kristy Loke +5
Cooperation between the United States and China, the world's leading artificial intelligence (AI) powers, is crucial for effective global AI governance and responsible AI developme…
Thinking beyond the anthropomorphic paradigm benefits LLM research
Lujain Ibrahim, Myra Cheng
Anthropomorphism, or the attribution of human traits to technology, is an automatic and unconscious response that occurs even in those with advanced technical expertise. In this po…
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar +7
The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for…