59 citations · 141 across the 54 of their papers we have counts for
3 papers · 2 filters
Reinforcement Learning with Verifiable Rewards for Small Search Agents
Gaurisankar Jayadas, Aske Plaat, Álvaro Serra-Gómez +1
Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is le…
Evaluation Metrics for Safe Reinforcement Learning
Lindsay Spoor, Aske Plaat, Thomas Moerland
Safe reinforcement learning (RL) is commonly formalized as a Constrained Markov Decision Process (CMDP), in which an agent maximizes expected reward while keeping its expected cumu…
Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring
Po-Chin Chang, Nicholas Hogan, Aske Plaat +1
LLMs can personalize education, although current static-prompt tutoring systems struggle to adapt to diverse academic disciplines. We develop and test a system with subject-aware p…