Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring
Usman Anwar, Julianna Piskorz, David D. Baek +6
Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to de…
cs.AI2025
Scaling Laws For Scalable Oversight
Joshua Engels, David D. Baek, Subhash Kantamneni +1
Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is s…