activity
20242026
collaborators

12 papers

cs.LG2026

Parameter Exploration for RLVR via Variational Learning

Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych

Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcem…

cs.LG2026

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

Hugo Monzón Maldonado, Nico Daheim, Thomas Möllenhoff +2

Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute. To reduce costs, recent works have used existing model m…

cs.LG2026

SVRG and Beyond via Posterior Correction

Nico Daheim, Thomas Möllenhoff, Ming Liang Ang +1

Stochastic Variance Reduced Gradient (SVRG) and its variants aim to speed-up training by using gradient corrections. Originally proposed over a decade ago, these methods have never…

cs.CL2026

'Your AI Text is not Mine': Redefining and Evaluating AI-generated Text Detection under Realistic Assumptions

Nils Dycke, Marina Sakharova, Nico Daheim +1

Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detection literature on what constitut…

cs.CL2025

From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning

David Dinucu-Jianu, Jakub Macina, Nico Daheim +3

Large language models (LLMs) can transform education, but their optimization for direct question-answering often undermines effective pedagogy which requires strategically withhold…

cs.CL2025

MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors

Jakub Macina, Nico Daheim, Ido Hakimi +3

Evaluating the pedagogical capabilities of AI-based tutoring models is critical for making guided progress in the field. Yet, we lack a reliable, easy-to-use, and simple-to-run eva…