activity
20172021
most citedPhonemic and Graphemic Multilingual CTC Based Speech Recognition

4 citations · 7 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2021

Instant One-Shot Word-Learning for Context-Specific Neural Sequence-to-Sequence Speech Recognition

Christian Huber, Juan Hussain, Sebastian Stüker +1

Neural sequence-to-sequence systems deliver state-of-the-art performance for automatic speech recognition (ASR). When using appropriate modeling units, e.g., byte-pair encoded char…

cs.CL2018

Neural Language Codes for Multilingual Acoustic Models

Markus Müller, Sebastian Stüker, Alex Waibel

Multilingual Speech Recognition is one of the most costly AI problems, because each language (7,000+) and even different accents require their own acoustic models to obtain best re…

cs.CL2018

Self-Attentional Acoustic Models

Matthias Sperber, Jan Niehues, Graham Neubig +2

Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising resul…

cs.CL2018

Linguistic unit discovery from multi-modal inputs in unwritten languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

Odette Scharenborg, Laurent Besacier, Alan Black +16

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and word…

cs.CL2017

Comparison of Decoding Strategies for CTC Acoustic Models

Thomas Zenkel, Ramon Sanabria, Florian Metze +4

Connectionist Temporal Classification has recently attracted a lot of interest as it offers an elegant approach to building acoustic models (AMs) for speech recognition. The CTC lo…

cs.CL20171 cited

Yeah, Right, Uh-Huh: A Deep Learning Backchannel Predictor

Robin Ruede, Markus Müller, Sebastian Stüker +1

Using supporting backchannel (BC) cues can make human-computer interaction more social. BCs provide a feedback from the listener to the speaker indicating to the speaker that he is…