1 citations · 1 across the 3 of their papers we have counts for
19 papers
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel +22
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchma…
Conversational AI increases political knowledge as effectively as self-directed internet search
Lennart Luettgau, Hannah Rose Kirk, Kobi Hackenburg +6
Conversational AI systems are increasingly being used in place of traditional search engines to help users complete information-seeking tasks. This has raised concerns in the polit…
AI systems out-persuade expert humans
Kobi Hackenburg, Caroline Wagner, Luke Hewitt +5
Many societal decisions are settled by contests of persuasion. Conversational AI is a powerful new entrant in these contests, but whether it can out-persuade skilled and highly inc…
RealityTest: How People Probe AI Identity and Whether Models Disclose It
Anna Gausen, Sarenne Wallbridge, Bessie O'Dell +2
AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mounting regulatory attention…
Post-training makes large language models less human-like
Marcel Binz, Elif Akata, Abdullah Almaatouq +76
Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, w…
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
Hannah Rose Kirk, Liu Leqi, Fanzhi Zeng +4
Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simu…