SpeechBrain: A General-Purpose Speech Toolkit
arXiv:2106.04624
Abstract
SpeechBrain is an open-source and all-in-one speech toolkit. It is designed to facilitate the research and development of neural speech processing technologies by being simple, flexible, user-friendly, and well-documented. This paper describes the core architecture designed to support several tasks of common interest, allowing users to naturally conceive, compare and share novel speech processing pipelines. SpeechBrain achieves competitive or state-of-the-art performance in a wide range of speech benchmarks. It also provides training recipes, pretrained models, and inference scripts for popular speech datasets, as well as tutorials which allow anyone with basic Python proficiency to familiarize themselves with speech technologies.
Preprint
References in corpus (11)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Sequence Transduction with Recurrent Neural Networks
- fastai: A Layered API for Deep Learning
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
- LibriMix: An Open-Source Dataset for Generalizable Speech Separation
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- BUT System Description to VoxCeleb Speaker Recognition Challenge 2019
- ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
Cited by in corpus (5)
- The SpeakIn System for VoxCeleb Speaker Recognition Challange 2021
- The ByteDance Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2021
- ASR4REAL: An extended benchmark for speech models
- Soundata: A Python library for reproducible use of audio datasets
- XMUSPEECH System for VoxCeleb Speaker Recognition Challenge 2021