4 papers
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models
Andres Algaba, Francesca Carlon, Lynn Delcon +3
Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds e…
Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments
Arne Vanhoyweghen, Vincent Holst, Melika Mobini +9
Providing timely and individualised feedback on handwritten student work is highly beneficial for learning but difficult to achieve at scale. This challenge has become more pressin…
Ergodicity in reinforcement learning
Dominik Baumann, Erfaun Noorani, Arsenii Mustafin +5
In reinforcement learning, we typically aim to optimize the expected value of the sum of rewards an agent collects over a trajectory. However, if the process generating these rewar…
Model-Agnostic Solutions for Deep Reinforcement Learning in Non-Ergodic Contexts
Bert Verbruggen, Arne Vanhoyweghen, Vincent Ginis
Reinforcement Learning (RL) remains a central optimisation framework in machine learning. Although RL agents can converge to optimal solutions, the definition of ``optimality'' dep…