156 citations
- Amazon (United States)US12 papers
- Massachusetts Institute of TechnologyUS7 papers
- Carnegie Mellon UniversityUS6 papers
- Google (United States)US5 papers
- Johns Hopkins UniversityUS5 papers
- Meta (Israel)IL5 papers
- California Southern UniversityUS4 papers
- Microsoft Research (United Kingdom)GB4 papers
- Toyota Technological Institute at ChicagoUS4 papers
- University of California, Los AngelesUS4 papers
- University of Southern CaliforniaUS4 papers
- National Yang Ming Chiao Tung UniversityTW3 papers
13 papers · 2 filters
Streaming Multi-speaker ASR with RNN-T
Ilya Sklyar, Anna Piunova, Yulan Liu
Recent research shows end-to-end ASR systems can recognize overlapped speech from multiple speakers. However, all published works have assumed no latency constraints during inferen…
Improving Device Directedness Classification of Utterances with Semantic Lexical Features
Kellen Gillespie, Ioannis C. Konstantakopoulos, Xingzhi Guo +2
User interactions with personal assistants like Alexa, Google Home and Siri are typically initiated by a wake term or wakeword. Several personal assistants feature "follow-up" mode…
Far-Field Automatic Speech Recognition
Reinhold Haeb-Umbach, Jahn Heymann, Lukas Drude +3
The machine recognition of speech spoken at a distance from the microphones, known as far-field automatic speech recognition (ASR), has received a significant increase of attention…
Speech Sentiment and Customer Satisfaction Estimation in Socialbot Conversations
Yelin Kim, Joshua Levy, Yang Liu
For an interactive agent, such as task-oriented spoken dialog systems or chatbots, measuring and adapting to Customer Satisfaction (CSAT) is critical in order to understand user pe…
Generating Music with a Self-Correcting Non-Chronological Autoregressive Model
Wayne Chi, Prachi Kumar, Suri Yaddanapudi +2
We describe a novel approach for generating music using a self-correcting, non-chronological, autoregressive model. We represent music as a sequence of edit events, each of which d…
A Joint Framework for Audio Tagging and Weakly Supervised Acoustic Event Detection Using DenseNet with Global Average Pooling
Chieh-Chi Kao, Bowen Shi, Ming Sun +1
This paper proposes a network architecture mainly designed for audio tagging, which can also be used for weakly supervised acoustic event detection (AED). The proposed network cons…