59 citations · 110 across the 49 of their papers we have counts for
8 papers · 1 filter
Evaluation Metrics for Safe Reinforcement Learning
Lindsay Spoor, Aske Plaat, Thomas Moerland
Safe reinforcement learning (RL) is commonly formalized as a Constrained Markov Decision Process (CMDP), in which an agent maximizes expected reward while keeping its expected cumu…
Silent Metronome: Rhythmic Grounding for Live Music Accompaniment
Kevin Bretz, Derya Soydaner, Aske Plaat
Live accompaniment models generate music for an incoming audio stream, committing to each output frame before hearing what comes next. In this strictly causal setting the model mus…
Detecting and Repairing Hallucinations in Retrieval-Augmented Generation
Sai Krishna Reddy Mulakkayala, Niki van Stein, Aske Plaat
Language models increasingly answer questions by consulting retrieved documents rather than memory alone, a design now common in search assistants and enterprise knowledge tools. G…
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18
Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…
Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring
Po-Chin Chang, Nicholas Hogan, Aske Plaat +1
LLMs can personalize education, although current static-prompt tutoring systems struggle to adapt to diverse academic disciplines. We develop and test a system with subject-aware p…
Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition
Po-Kai Chen, Aske Plaat, Niki van Stein
Mechanistic interpretability of transformers requires identifying not just which components matter but how they compose into the computational route that produced a prediction. Bot…