16 citations · 17 across the 6 of their papers we have counts for
9 papers
Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
Carel van Niekerk, Renato Vukovic, Benjamin Matthias Ruppik +2
Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks. We present Reinforcement Learning from…
Emotionally Intelligent Task-oriented Dialogue Systems: Architecture, Representation, and Optimisation
Shutong Feng, Hsien-chin Lin, Nurul Lubis +5
Task-oriented dialogue (ToD) systems are designed to help users achieve specific goals through natural language interaction. While recent advances in large language models (LLMs) h…
What Does The User Want? Information Gain for Hierarchical Dialogue Policy Optimisation
Christian Geishauser, Songbo Hu, Hsien-chin Lin +5
The dialogue management component of a task-oriented dialogue system is typically optimised via reinforcement learning (RL). Optimisation via RL is highly susceptible to sample ine…
Uncertainty Measures in Neural Belief Tracking and the Effects on Dialogue Policy Performance
Carel van Niekerk, Andrey Malinin, Christian Geishauser +5
The ability to identify and resolve uncertainty is crucial for the robustness of a dialogue system. Indeed, this has been confirmed empirically on systems that utilise Bayesian app…
Domain-independent User Simulation with Transformers for Task-oriented Dialogue Systems
Hsien-chin Lin, Nurul Lubis, Songbo Hu +5
Dialogue policy optimisation via reinforcement learning requires a large number of training interactions, which makes learning with real users time consuming and expensive. Many se…
Out-of-Task Training for Dialog State Tracking Models
Michael Heck, Carel van Niekerk, Nurul Lubis +4
Dialog state tracking (DST) suffers from severe data sparsity. While many natural language processing (NLP) tasks benefit from transfer learning and multi-task learning, in dialog…