5 citations · 8 across the 2 of their papers we have counts for
6 papers
Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems
Anish Acharya, Suranjit Adhikari, Sanchit Agarwal +28
Traditional goal-oriented dialogue systems rely on various components such as natural language understanding, dialogue state tracking, policy learning and response generation. Trai…
Realizing Petabyte Scale Acoustic Modeling
Sree Hari Krishnan Parthasarathi, Nitin Sivakrishnan, Pranav Ladkat +1
Large scale machine learning (ML) systems such as the Alexa automatic speech recognition (ASR) system continue to improve with increasing amounts of manually transcribed training d…
Lessons from Building Acoustic Models with a Million Hours of Speech
Sree Hari Krishnan Parthasarathi, Nikko Strom
This is a report of our lessons learned building acoustic models from 1 Million hours of unlabeled speech, while labeled speech is restricted to 7,000 hours. We employ student/teac…
Comprehensive evaluation of statistical speech waveform synthesis
Thomas Merritt, Bartosz Putrycz, Adam Nadolski +10
Statistical TTS systems that directly predict the speech waveform have recently reported improvements in synthesis quality. This investigation evaluates Amazon's statistical speech…
Data Augmentation for Robust Keyword Spotting under Playback Interference
Anirudh Raju, Sankaran Panchapagesan, Xing Liu +2
Accurate on-device keyword spotting (KWS) with low false accept and false reject rate is crucial to customer experience for far-field voice control of conversational agents. It is…
Max-Pooling Loss Training of Long Short-Term Memory Networks for Small-Footprint Keyword Spotting
Ming Sun, Anirudh Raju, George Tucker +6
We propose a max-pooling based loss function for training Long Short-Term Memory (LSTM) networks for small-footprint keyword spotting (KWS), with low CPU, memory, and latency requi…