3 papers
cs.LG2026
Adaptive Reinforcement Learning for Unobservable Random Delays
John Wikman, Alexandre Proutiere, David Broman
In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that…
cs.AI2026
Advantage-Guided Diffusion for Model-Based Reinforcement Learning
Daniele Foffano, Arvid Eriksson, David Broman +2
Model-based reinforcement learning (MBRL) with autoregressive world models suffers from compounding errors, whereas diffusion world models mitigate this by generating trajectory se…
cs.AI2024
Learning Formal Mathematics From Intrinsic Motivation
Gabriel Poesia, David Broman, Nick Haber +1
How did humanity coax mathematics from the aether? We explore the Platonic view that mathematics can be discovered from its axioms - a game of conjecture and proof. We describe Min…