activity
20242026
collaborators

10 papers

cs.CY2026

Europe and the Geopolitics of AGI: The Need for a Preparedness Plan

Maximilian Negele, Daan Juijn, Afek Shamir +8

Artificial general intelligence (AGI)--defined here as AI systems that match or exceed humans at most economically useful cognitive work--has moved from speculation to the centre o…

cs.AI2026

Seven simple steps for log analysis in AI systems

Magda Dubois, Ekin Zorer, Maia Hamin +17

AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess…

cs.CL2026

No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes

Iván Vicente Moreno Cencerrado, Arnau Padrés Masdemont, Anton Gonzalvez Hawthorne +2

Do large language models (LLMs) anticipate when they will answer correctly? To study this, we extract activations after a question is read but before any tokens are generated, and…

cs.LG2026

Capabilities Ain't All You Need: Measuring Propensities in AI

Daniel Romero-Alvarado, Fernando Martínez-Plumed, Lorenzo Pacchiardi +11

AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the te…

cs.CY2026

Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies

Miles Brundage, Noemi Dreksler, Aidan Homewood +45

We outline a vision for frontier AI auditing, which we define as rigorous third-party verification of frontier AI developers' safety and security claims, and evaluation of their sy…

cs.AI2025

Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents

Irene Testini, José Hernández-Orallo, Lorenzo Pacchiardi

Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data scie…