3 papers
cs.SE2026
Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
Yanuo Ma, Ben Kereopa-Yorke, Ben Schultz
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity problems, and a passing score may no…
cs.CR2026
Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
Ben Kereopa-Yorke, Guillermo Diaz, Holly Wright +3
We define Oracle Poisoning, an attack class in which an adversary corrupts a structured knowledge graph that AI agents query at runtime via tool-use protocols, causing incorrect co…
cs.HC2025
Engineering Trust, Creating Vulnerability: A Socio-Technical Analysis of AI Interface Design
Ben Kereopa-Yorke
This paper examines how distinct cultures of AI interdisciplinarity emerge through interface design, revealing the formation of new disciplinary cultures at these intersections. Th…