6 papers
Fast Adversarial Attacks with Gradient Prediction
Kamil Ciosek, Aleksandr V. Petrov, Nicolò Felicioni +1
Generating adversarial examples at scale is a core primitive for robustness evaluation, adversarial training, and red-teaming, yet even "fast" attacks such as FGSM remain throughpu…
Position: agentic AI orchestration should be Bayes-consistent
Theodore Papamarkou, Pierre Alquier, Matthias Bauer +27
LLMs excel at predictive tasks and complex reasoning tasks, but many high-value deployments rely on decisions under uncertainty, for example, which tool to call, which expert to co…
A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks
Sergio Calvo-Ordoñez, Jonathan Plenk, Richard Bergna +4
Performing gradient descent in a wide neural network is equivalent to computing the posterior mean of a Gaussian Process with the Neural Tangent Kernel (NTK-GP), for a specific pri…
Mapping Faithful Reasoning in Language Models
Jiazheng Li, Andreas Damianou, J Rosser +2
Chain-of-thought (CoT) traces promise transparency for reasoning language models, but prior work shows they are not always faithful reflections of internal computation. This raises…
Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models
Konstantina Palla, José Luis Redondo GarcÃa, Claudia Hauff +5
Content moderation plays a critical role in shaping safe and inclusive online environments, balancing platform standards, user expectations, and regulatory frameworks. Traditionall…
Enhancing Content Moderation with Culturally-Aware Models
Alex J. Chan, José Luis Redondo GarcÃa, Fabrizio Silvestri +2
Content moderation on a global scale must navigate a complex array of local cultural distinctions, which can hinder effective enforcement. While global policies aim for consistency…