2 citations · 2 across the 10 of their papers we have counts for
10 papers
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
Mohan Li, Cong-Thanh Do, Simon Keizer +3
Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this p…
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
Cong-Thanh Do, Shuhei Imai, Rama Doddipatla +1
This paper investigates the use of unsupervised text-to-speech synthesis (TTS) as a data augmentation method to improve accented speech recognition. TTS systems are trained with a…
Semantic Map-based Generation of Navigation Instructions
Chengzu Li, Chao Zhang, Simone Teufel +2
We are interested in the generation of navigation instructions, either in their own right or as training material for robotic navigation task. In this paper, we propose a new appro…
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partiall…
Evaluating Large Language Models for Document-grounded Response Generation in Information-Seeking Dialogues
Norbert Braunschweiler, Rama Doddipatla, Simon Keizer +1
In this paper, we investigate the use of large language models (LLMs) like ChatGPT for document-grounded response generation in the context of information-seeking dialogues. For ev…
Adversarial learning of neural user simulators for dialogue policy optimisation
Simon Keizer, Caroline Dockes, Norbert Braunschweiler +2
Reinforcement learning based dialogue policies are typically trained in interaction with a user simulator. To obtain an effective and robust policy, this simulator should generate…