Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Inverting the Bellman Equation: From -Values to World Models
Alistair Letcher, Mattie Fellows, Alexander D. Goldie +3
Model-based and model-free reinforcement learning are traditionally viewed as separate paradigms: instead of learning a model of the transition kernel , model-free agents typica…
cs.LG2026
Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
Minsu Kim, Jean-Pierre Falet, Oliver E. Richardson +5
Chain-of-Thought (CoT) reasoning has advanced the capabilities and transparency of language models (LMs); however, reasoning chains can contain inaccurate statements that reduce pe…
cs.LG2025
Learning with Confidence
Oliver Ethan Richardson
We characterize a notion of confidence that arises in learning or updating beliefs: the amount of trust one has in incoming information and its impact on the belief state. This lea…