activity
20172022
most citedDo End-to-End Speech Recognition Models Care About Context?

9 citations · 14 across the 5 of their papers we have counts for

collaborators

7 papers

eess.AS20225 cited

A Brief Overview of Unsupervised Neural Speech Representation Learning

Lasse Borgholt, Jakob Drachmann Havtorn, Joakim Edin +2

Unsupervised representation learning for speech processing has matured greatly in the last few years. Work in computer vision and natural language processing has paved the way, but…

eess.AS2022

Benchmarking Generative Latent Variable Models for Speech

Jakob D. Havtorn, Lasse Borgholt, Søren Hauberg +2

Stochastic latent variable models (LVMs) achieve state-of-the-art performance on natural image generation but are still inferior to deterministic models on speech. In this paper, w…

eess.AS20219 cited

Do End-to-End Speech Recognition Models Care About Context?

Lasse Borgholt, Jakob Drachmann Havtorn, Željko Agić +3

The two most common paradigms for end-to-end speech recognition are connectionist temporal classification (CTC) and attention-based encoder-decoder (AED) models. It has been argued…

eess.AS2021

On Scaling Contrastive Representations for Low-Resource Speech Recognition

Lasse Borgholt, Tycho Max Sylvester Tax, Jakob Drachmann Havtorn +2

Recent advances in self-supervised learning through contrastive training have shown that it is possible to learn a competitive speech recognition system with as little as 10 minute…

cs.CL2020

MultiQT: Multimodal Learning for Real-Time Question Tracking in Speech

Jakob D. Havtorn, Jan Latko, Joakim Edin +6

We address a challenging and practical task of labeling questions in speech in real time during telephone calls to emergency medical services in English, which embeds within a broa…

cs.CL2018

On the Inductive Bias of Word-Character-Level Multi-Task Learning for Speech Recognition

Jan Kremer, Lasse Borgholt, Lars Maaløe

End-to-end automatic speech recognition (ASR) commonly transcribes audio signals into sequences of characters while its performance is evaluated by measuring the word-error rate (W…