activity
20222024
most citedMultilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study

1 citations · 1 across the 7 of their papers we have counts for

collaborators

6 papers

cs.CL2023

Text Injection for Capitalization and Turn-Taking Prediction in Speech Models

Shaan Bijwadia, Shuo-yiin Chang, Weiran Wang +3

Text injection for automatic speech recognition (ASR), wherein unpaired text-only data is used to supplement paired audio-text data, has shown promising improvements for word error…

cs.CL2023

Semantic Segmentation with Bidirectional Language Models Improves Long-form ASR

W. Ronny Huang, Hao Zhang, Shankar Kumar +2

We propose a method of segmenting long-form speech by separating semantically complete sentences within the utterance. This prevents the ASR decoder from needlessly processing fara…

eess.AS2023

UML: A Universal Monolingual Output Layer for Multilingual ASR

Chao Zhang, Bo Li, Tara N. Sainath +2

Word-piece models (WPMs) are commonly used subword units in state-of-the-art end-to-end automatic speech recognition (ASR) systems. For multilingual ASR, due to the differences in…

eess.AS2022

A Language Agnostic Multilingual Streaming On-Device ASR System

Bo Li, Tara N. Sainath, Ruoming Pang +9

On-device end-to-end (E2E) models have shown improvements over a conventional model on English Voice Search tasks in both quality and latency. E2E models have also shown promising…

cs.CL2022

Streaming Intended Query Detection using E2E Modeling for Continued Conversation

Shuo-yiin Chang, Guru Prakash, Zelin Wu +7

In voice-enabled applications, a predetermined hotword isusually used to activate a device in order to attend to the query.However, speaking queries followed by a hotword each time…

cs.CL2022

Turn-Taking Prediction for Natural Conversational Speech

Shuo-yiin Chang, Bo Li, Tara N. Sainath +4

While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice qu…