4 papers · 1 filter
Can adversarial attacks by large language models be attributed?
Manuel Cebrian, Andres Abeliuk, Jan Arne Telle
Attributing outputs from Large Language Models (LLMs) in adversarial settings-such as cyberattacks and disinformation campaigns-presents significant challenges that are likely to g…
Supervision policies can shape long-term risk management in general-purpose AI models
Manuel Cebrian, Emilia Gomez, David Fernandez Llorca
The rapid proliferation and deployment of General-Purpose AI (GPAI) models, including large language models (LLMs), present unprecedented challenges for AI supervisory entities. We…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando MartÃnez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
Conversational Complexity for Assessing Risk in Large Language Models
John Burden, Manuel Cebrian, Jose Hernandez-Orallo
Large Language Models (LLMs) present a dual-use dilemma: they enable beneficial applications while harboring potential for harm, particularly through conversational interactions. D…