4 papers · 1 filter
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
Language models recognize dropout and Gaussian noise applied to their activations
Damiano Fornasiere, Mirko Bronzi, Spencer Kitts +3
We provide evidence that language models can detect, localize and, to a certain degree, verbalize the difference between perturbations applied to their activations. More precisely,…
Local Inconsistency Resolution: The Interplay between Attention and Control in Probabilistic Models
Oliver E. Richardson, Mandana Samiei, Mehran Shakerinava +4
We present a generic algorithm for learning and approximate inference with an intuitive epistemic interpretation: iteratively focus on a subset of the model and resolve inconsisten…
Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
Yoshua Bengio, Michael Cohen, Damiano Fornasiere +10
The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans…