9 citations · 14 across the 5 of their papers we have counts for
7 papers
A Brief Overview of Unsupervised Neural Speech Representation Learning
Lasse Borgholt, Jakob Drachmann Havtorn, Joakim Edin +2
Unsupervised representation learning for speech processing has matured greatly in the last few years. Work in computer vision and natural language processing has paved the way, but…
Benchmarking Generative Latent Variable Models for Speech
Jakob D. Havtorn, Lasse Borgholt, Søren Hauberg +2
Stochastic latent variable models (LVMs) achieve state-of-the-art performance on natural image generation but are still inferior to deterministic models on speech. In this paper, w…
Do End-to-End Speech Recognition Models Care About Context?
Lasse Borgholt, Jakob Drachmann Havtorn, Željko Agić +3
The two most common paradigms for end-to-end speech recognition are connectionist temporal classification (CTC) and attention-based encoder-decoder (AED) models. It has been argued…
On Scaling Contrastive Representations for Low-Resource Speech Recognition
Lasse Borgholt, Tycho Max Sylvester Tax, Jakob Drachmann Havtorn +2
Recent advances in self-supervised learning through contrastive training have shown that it is possible to learn a competitive speech recognition system with as little as 10 minute…
MultiQT: Multimodal Learning for Real-Time Question Tracking in Speech
Jakob D. Havtorn, Jan Latko, Joakim Edin +6
We address a challenging and practical task of labeling questions in speech in real time during telephone calls to emergency medical services in English, which embeds within a broa…
On the Inductive Bias of Word-Character-Level Multi-Task Learning for Speech Recognition
Jan Kremer, Lasse Borgholt, Lars Maaløe
End-to-end automatic speech recognition (ASR) commonly transcribes audio signals into sequences of characters while its performance is evaluated by measuring the word-error rate (W…