Sequence-to-Sequence Neural Net Models for Grapheme-to-Phoneme Conversion
arXiv:1506.00196
Abstract
Sequence-to-sequence translation methods based on generation with a side-conditioned language model have recently shown promising results in several tasks. In machine translation, models conditioned on source side words have been used to produce target-language text, and in image captioning, models conditioned images have been used to generate caption text. Past work with this approach has focused on large vocabulary tasks, and measured quality in terms of BLEU. In this paper, we explore the applicability of such models to the qualitatively different grapheme-to-phoneme task. Here, the input and output side vocabularies are small, plain n-gram models do well, and credit is only given when the output is exactly correct. We find that the simple side-conditioned generation approach is able to rival the state-of-the-art, and we are able to significantly advance the stat-of-the-art with bi-directional long short-term memory (LSTM) neural networks that use the same alignment information that is used in conventional approaches.
Published in INTERSPEECH 2015, Dresden, Germany
References in corpus (4)
Cited by in corpus (16)
- Deep Voice: Real-time Neural Text-to-Speech
- A Survey on Neural Speech Synthesis
- Attention with Intention for a Neural Network Conversation Model
- Sequence-to-sequence neural network models for transliteration
- Transformer based Grapheme-to-Phoneme Conversion
- Understanding and Mitigating the Security Risks of Voice-Controlled Third-Party Skills on Amazon Alexa and Google Home
- Design Challenges in Named Entity Transliteration
- Leveraging Sentence-level Information with Encoder LSTM for Semantic Slot Filling
- Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion
- Still not there? Comparing Traditional Sequence-to-Sequence Models to Encoder-Decoder Neural Networks on Monotone String Translation Tasks
- Natural Language Generation with Neural Variational Models
- Delhi air quality prediction using LSTM deep learning models with a focus on COVID-19 lockdown
- Simple Models for Word Formation in English Slang
- Massively Multilingual Neural Grapheme-to-Phoneme Conversion
- Jointly Learning to Align and Convert Graphemes to Phonemes with Neural Attention Models
- TFW, DamnGina, Juvie, and Hotsie-Totsie: On the Linguistic and Social Aspects of Internet Slang