1 citations · 1 across the 3 of their papers we have counts for
5 papers
RealityTest: How People Probe AI Identity and Whether Models Disclose It
Anna Gausen, Sarenne Wallbridge, Bessie O'Dell +2
AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mounting regulatory attention…
A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime Scenarios
Kimberly T. Mai, Anna Gausen, Magda Dubois +5
AI is increasingly being used to assist fraud and cybercrime. However, it is unclear the extent to which current large language models can provide useful information for complex cr…
Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
Anna Gausen, Sarenne Wallbridge, Hannah Rose Kirk +2
As conversational AI systems become more realistic and widely deployed, users are increasingly uncertain about whether they are interacting with a human or an AI system. When AI id…
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…
A Framework for Exploring the Consequences of AI-Mediated Enterprise Knowledge Access and Identifying Risks to Workers
Anna Gausen, Bhaskar Mitra, Siân Lindley
Organisations generate vast amounts of information, which has resulted in a long-term research effort into knowledge access systems for enterprise settings. Recent developments in…