collaborators

9 papers

cs.CC2026

How to Verify Consistency of Probabilistic Claims

Orr Paradise, Oliver Richardson, Yoshua Bengio +1

When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of intere…

cs.AI2026

Safety from Honesty in a Disinterested AI Predictor

Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…

cs.LG2026

Inverting the Bellman Equation: From -Values to World Models

Alistair Letcher, Mattie Fellows, Alexander D. Goldie +3

Model-based and model-free reinforcement learning are traditionally viewed as separate paradigms: instead of learning a model of the transition kernel , model-free agents typica…

cs.AI2026

Language models recognize dropout and Gaussian noise applied to their activations

Damiano Fornasiere, Mirko Bronzi, Spencer Kitts +3

We provide evidence that language models can detect, localize and, to a certain degree, verbalize the difference between perturbations applied to their activations. More precisely,…

cs.AI2026

Local Inconsistency Resolution: The Interplay between Attention and Control in Probabilistic Models

Oliver E. Richardson, Mandana Samiei, Mehran Shakerinava +4

We present a generic algorithm for learning and approximate inference with an intuitive epistemic interpretation: iteratively focus on a subset of the model and resolve inconsisten…

cs.LG2026

Latent Veracity Inference for Identifying Errors in Stepwise Reasoning

Minsu Kim, Jean-Pierre Falet, Oliver E. Richardson +5

Chain-of-Thought (CoT) reasoning has advanced the capabilities and transparency of language models (LMs); however, reasoning chains can contain inaccurate statements that reduce pe…