Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models
Andres Algaba, Francesca Carlon, Lynn Delcon +3
Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds e…
cs.LG2026
Ergodicity in reinforcement learning
Dominik Baumann, Erfaun Noorani, Arsenii Mustafin +5
In reinforcement learning, we typically aim to optimize the expected value of the sum of rewards an agent collects over a trajectory. However, if the process generating these rewar…
cs.LG2026
Model-Agnostic Solutions for Deep Reinforcement Learning in Non-Ergodic Contexts
Bert Verbruggen, Arne Vanhoyweghen, Vincent Ginis
Reinforcement Learning (RL) remains a central optimisation framework in machine learning. Although RL agents can converge to optimal solutions, the definition of ``optimality'' dep…