4 papers
Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
Carel van Niekerk, Renato Vukovic, Benjamin Matthias Ruppik +2
Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks. We present Reinforcement Learning from…
Emotionally Intelligent Task-oriented Dialogue Systems: Architecture, Representation, and Optimisation
Shutong Feng, Hsien-chin Lin, Nurul Lubis +5
Task-oriented dialogue (ToD) systems are designed to help users achieve specific goals through natural language interaction. While recent advances in large language models (LLMs) h…
Text-to-SQL Task-oriented Dialogue Ontology Construction
Renato Vukovic, Carel van Niekerk, Michael Heck +5
Large language models (LLMs) are widely used as general-purpose knowledge sources, but they rely on parametric knowledge, limiting explainability and trustworthiness. In task-orien…
Learning from Noisy Labels via Self-Taught On-the-Fly Meta Loss Rescaling
Michael Heck, Christian Geishauser, Nurul Lubis +6
Correct labels are indispensable for training effective machine learning models. However, creating high-quality labels is expensive, and even professionally labeled data contains e…