1 citations · 1 across the 7 of their papers we have counts for
6 papers
Text Injection for Capitalization and Turn-Taking Prediction in Speech Models
Shaan Bijwadia, Shuo-yiin Chang, Weiran Wang +3
Text injection for automatic speech recognition (ASR), wherein unpaired text-only data is used to supplement paired audio-text data, has shown promising improvements for word error…
Semantic Segmentation with Bidirectional Language Models Improves Long-form ASR
W. Ronny Huang, Hao Zhang, Shankar Kumar +2
We propose a method of segmenting long-form speech by separating semantically complete sentences within the utterance. This prevents the ASR decoder from needlessly processing fara…
UML: A Universal Monolingual Output Layer for Multilingual ASR
Chao Zhang, Bo Li, Tara N. Sainath +2
Word-piece models (WPMs) are commonly used subword units in state-of-the-art end-to-end automatic speech recognition (ASR) systems. For multilingual ASR, due to the differences in…
A Language Agnostic Multilingual Streaming On-Device ASR System
Bo Li, Tara N. Sainath, Ruoming Pang +9
On-device end-to-end (E2E) models have shown improvements over a conventional model on English Voice Search tasks in both quality and latency. E2E models have also shown promising…
Streaming Intended Query Detection using E2E Modeling for Continued Conversation
Shuo-yiin Chang, Guru Prakash, Zelin Wu +7
In voice-enabled applications, a predetermined hotword isusually used to activate a device in order to attend to the query.However, speaking queries followed by a hotword each time…
Turn-Taking Prediction for Natural Conversational Speech
Shuo-yiin Chang, Bo Li, Tara N. Sainath +4
While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice qu…