Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring
Usman Anwar, Julianna Piskorz, David D. Baek +6
Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to de…
cs.AI2025
The Reality of AI and Biorisk
Aidan Peppin, Anka Reuel, Stephen Casper +10
To accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or…
cs.AI2024
IDs for AI Systems
Alan Chan, Noam Kolt, Peter Wills +7
AI systems are increasingly pervasive, yet information needed to decide whether and how to engage with them may not exist or be accessible. A user may not be able to verify whether…