4 citations · 4 across the 2 of their papers we have counts for
3 papers
cs.AI2026
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
Rebecca Ansell, Autumn Toney-Wails
Deducing whodunit proves challenging for LLM agents. In this paper, we implement a text-based multi-agent version of the classic board game Clue as a rule-based testbed for evaluat…
cs.CL2025
Certain but not Probable? Differentiating Certainty from Probability in LLM Token Outputs for Probabilistic Scenarios
Autumn Toney-Wails, Ryan Wails
Reliable uncertainty quantification (UQ) is essential for ensuring trustworthy downstream use of large language models, especially when they are deployed in decision-support and ot…
cs.CL2024★ 4 cited
AI on AI: Exploring the Utility of GPT as an Expert Annotator of AI Publications
Autumn Toney-Wails, Christian Schoeberl, James Dunham
Identifying scientific publications that are within a dynamic field of research often requires costly annotation by subject-matter experts. Resources like widely-accepted classific…