4 papers
Emergent evaluation hubs in a decentralizing large language model ecosystem
Manuel Cebrian, Tomomi Kito, Raul Castro Fernandez
Large language models are proliferating, and so are the benchmarks that serve as their common yardsticks. We ask how the agglomeration patterns of these two layers compare: do they…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
Supervision policies can shape long-term risk management in general-purpose AI models
Manuel Cebrian, Emilia Gomez, David Fernandez Llorca
The rapid proliferation and deployment of General-Purpose AI (GPAI) models, including large language models (LLMs), present unprecedented challenges for AI supervisory entities. We…
Can adversarial attacks by large language models be attributed?
Manuel Cebrian, Andres Abeliuk, Jan Arne Telle
Attributing outputs from Large Language Models (LLMs) in adversarial settings-such as cyberattacks and disinformation campaigns-presents significant challenges that are likely to g…