3 papers
cs.CL2026
Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue
William Hager, Ishika Rathi, Masum Hasan +1
As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates…
cs.HC2024
GPT-4 is judged more human than humans in displaced and inverted Turing tests
Ishika Rathi, Sydney Taylor, Benjamin K. Bergen +1
Everyday AI detection requires differentiating between people and AI in informal, online conversations. In many cases, people will not interact directly with AI systems but instead…
cs.HC2024
People cannot distinguish GPT-4 from a human in a Turing test
Cameron R. Jones, Benjamin K. Bergen
We evaluated 3 systems (ELIZA, GPT-3.5 and GPT-4) in a randomized, controlled, and preregistered Turing test. Human participants had a 5 minute conversation with either a human or…