11 citations · 12 across the 9 of their papers we have counts for
6 papers · 1 filter
Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Matthew Le, Apoorv Vyas, Bowen Shi +8
Large-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists whic…
A Joint Framework for Audio Tagging and Weakly Supervised Acoustic Event Detection Using DenseNet with Global Average Pooling
Chieh-Chi Kao, Bowen Shi, Ming Sun +1
This paper proposes a network architecture mainly designed for audio tagging, which can also be used for weakly supervised acoustic event detection (AED). The proposed network cons…
Whole-Word Segmental Speech Recognition with Acoustic Word Embeddings
Bowen Shi, Shane Settle, Karen Livescu
Segmental models are sequence prediction models in which scores of hypotheses are based on entire variable-length segments of frames. We consider segmental models for whole-word ("…
Compression of Acoustic Event Detection Models With Quantized Distillation
Bowen Shi, Ming Sun, Chieh-Chi Kao +3
Acoustic Event Detection (AED), aiming at detecting categories of events based on audio signals, has found application in many intelligent systems. Recently deep neural network sig…
Compression of Acoustic Event Detection Models with Low-rank Matrix Factorization and Quantization Training
Bowen Shi, Ming Sun, Chieh-Chi Kao +3
In this paper, we present a compression approach based on the combination of low-rank matrix factorization and quantization training, to reduce complexity for neural network based…
Semi-supervised Acoustic Event Detection based on tri-training
Bowen Shi, Ming Sun, Chieh-Chi Kao +3
This paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, wh…