Publications (30)
Parameter Exploration for RLVR via Variational Learning
Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcem…
Uncertainty-Aware Generation and Decision-Making Under Ambiguity
Nico Daheim, Iryna Gurevych
With rapidly improving capabilities, Large Language Models (LLMs) are increasingly used in many complex real-world tasks. Beyond requiring in-depth knowledge and reasoning skills,…
Improving LoRA with Variational Learning
Bai Cong, Nico Daheim, Yuesong Shen +3
Bayesian methods have recently been used to improve LoRA finetuning and, although they improve calibration, their effect on other metrics (such as accuracy) is marginal and can som…
Controllable Factuality in Document-Grounded Dialog Systems Using a Noisy Channel Model
Nico Daheim, David Thulke, Christian Dugast +1
In this work, we present a model for document-grounded response generation in dialog that is decomposed into two components according to Bayes theorem. One component is a tradition…
Cascaded Span Extraction and Response Generation for Document-Grounded Dialog
Nico Daheim, David Thulke, Christian Dugast +1
This paper summarizes our entries to both subtasks of the first DialDoc shared task which focuses on the agent response prediction task in goal-oriented document-grounded dialogs.…
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems
Jakub Macina, Nico Daheim, Sankalan Pal Chowdhury +4
While automatic dialogue tutors hold great potential in making education personalized and more accessible, research on such systems has been hampered by a lack of sufficiently larg…