Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents
Wael Albayaydh, Rui Zhao, Ivan Flechais
Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons.…
cs.AI2026
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and behave differently under thos…
cs.AI2026
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction f…