9 papers
How to Verify Consistency of Probabilistic Claims
Orr Paradise, Oliver Richardson, Yoshua Bengio +1
When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of intere…
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
Inverting the Bellman Equation: From -Values to World Models
Alistair Letcher, Mattie Fellows, Alexander D. Goldie +3
Model-based and model-free reinforcement learning are traditionally viewed as separate paradigms: instead of learning a model of the transition kernel , model-free agents typica…
Language models recognize dropout and Gaussian noise applied to their activations
Damiano Fornasiere, Mirko Bronzi, Spencer Kitts +3
We provide evidence that language models can detect, localize and, to a certain degree, verbalize the difference between perturbations applied to their activations. More precisely,…
Local Inconsistency Resolution: The Interplay between Attention and Control in Probabilistic Models
Oliver E. Richardson, Mandana Samiei, Mehran Shakerinava +4
We present a generic algorithm for learning and approximate inference with an intuitive epistemic interpretation: iteratively focus on a subset of the model and resolve inconsisten…
Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
Minsu Kim, Jean-Pierre Falet, Oliver E. Richardson +5
Chain-of-Thought (CoT) reasoning has advanced the capabilities and transparency of language models (LMs); however, reasoning chains can contain inaccurate statements that reduce pe…