SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
arXiv:1904.08779 · doi:10.21437/Interspeech.2019-2680
Abstract
We present SpecAugment, a simple data augmentation method for speech recognition. SpecAugment is applied directly to the feature inputs of a neural network (i.e., filter bank coefficients). The augmentation policy consists of warping the features, masking blocks of frequency channels, and masking blocks of time steps. We apply SpecAugment on Listen, Attend and Spell networks for end-to-end speech recognition tasks. We achieve state-of-the-art performance on the LibriSpeech 960h and Swichboard 300h tasks, outperforming all prior work. On LibriSpeech, we achieve 6.8% WER on test-other without the use of a language model, and 5.8% WER with shallow fusion with a language model. This compares to the previous state-of-the-art hybrid system of 7.5% WER. For Switchboard, we achieve 7.2%/14.6% on the Switchboard/CallHome portion of the Hub5'00 test set without the use of a language model, and 6.8%/14.1% with shallow fusion, which compares to the previous state-of-the-art hybrid system at 8.3%/17.3% WER.
5 pages, 3 figures, 6 tables; v3: references added
References in corpus (3)
Cited by in corpus (618)
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Unsupervised Data Augmentation for Consistency Training
- ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
- Contrastive Representation Learning: A Framework and Review
- A Comparative Study on Transformer vs RNN in Speech Applications
- An Empirical Survey of Data Augmentation for Time Series Classification with Neural Networks
- SpeechBrain: A General-Purpose Speech Toolkit
- Conformer: Convolution-augmented Transformer for Speech Recognition
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- A Survey of Sound Source Localization with Deep Learning Methods
- Data Augmentation techniques in time series domain: A survey and taxonomy
- Multi-Attention-Network for Semantic Segmentation of Fine Resolution Remote Sensing Images
- A2-FPN for Semantic Segmentation of Fine-Resolution Remotely Sensed Images
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Attention Bottlenecks for Multimodal Fusion
- Transformer-based Acoustic Modeling for Hybrid Speech Recognition
- Improved Noisy Student Training for Automatic Speech Recognition
- Efficient Training of Audio Transformers with Patchout
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- Self-Supervised MultiModal Versatile Networks
- Language Modeling with Deep Transformers
- Land Cover Classification from Remote Sensing Images Based on Multi-Scale Fully Convolutional Network
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
- PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation
- Visual Speech Recognition for Multiple Languages in the Wild
- Deep Contextualized Acoustic Representations For Semi-Supervised Speech Recognition
- Keyword Transformer: A Self-Attention Model for Keyword Spotting
- MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers
- Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019
- AutoML-Zero: Evolving Machine Learning Algorithms From Scratch
- Perceiver: General Perception with Iterative Attention
- Streaming keyword spotting on mobile devices
- Fine-tuning wav2vec2 for speaker recognition
- Effectiveness of self-supervised pre-training for speech recognition
- ECAPA-TDNN Embeddings for Speaker Diarization
- Cough Against COVID: Evidence of COVID-19 Signature in Cough Sounds
- Recent Progress in the CUHK Dysarthric Speech Recognition System
- Data Augmenting Contrastive Learning of Speech Representations in the Time Domain
- fairseq S2T: Fast Speech-to-Text Modeling with fairseq
- Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition
- ERANNs: Efficient Residual Audio Neural Networks for Audio Pattern Recognition
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- Sound source detection, localization and classification using consecutive ensemble of CRNN models
- Training Speech Recognition Models with Federated Learning: A Quality/Cost Framework
- PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
- ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
- Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture
- Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
- Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification
- SALSA: Spatial Cue-Augmented Log-Spectrogram Features for Polyphonic Sound Event Localization and Detection
- Training Strategies for Improved Lip-reading
- The IDLAB VoxSRC-20 Submission: Large Margin Fine-Tuning and Quality-Aware Score Calibration in DNN Based Speaker Verification
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Visual-Tactile Cross-Modal Data Generation using Residue-Fusion GAN with Feature-Matching and Perceptual Losses
- You Only Hear Once: A YOLO-like Algorithm for Audio Segmentation and Sound Event Detection
- Affinity and Diversity: Quantifying Mechanisms of Data Augmentation
- Investigation of Data Augmentation Techniques for Disordered Speech Recognition
- Towards duration robust weakly supervised sound event detection
- A Two-Stage Approach to Device-Robust Acoustic Scene Classification
- Device-Robust Acoustic Scene Classification Based on Two-Stage Categorization and Data Augmentation
- Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
- Transfer Learning and SpecAugment applied to SSVEP Based BCI Classification
- Jasper: An End-to-End Convolutional Neural Acoustic Model
- CoVoST 2 and Massively Multilingual Speech-to-Text Translation
- Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- A New Training Pipeline for an Improved Neural Transducer
- Recent Developments on ESPnet Toolkit Boosted by Conformer
- Bridging the Modality Gap for Speech-to-Text Translation
- Earnings-21: A Practical Benchmark for ASR in the Wild
- On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition
- SALSA-Lite: A Fast and Effective Feature for Polyphonic Sound Event Localization and Detection with Microphone Arrays
- Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention
- Multiclass Language Identification using Deep Learning on Spectral Images of Audio Signals
- Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
- An Overview of Indian Spoken Language Recognition from Machine Learning Perspective
- Successes and critical failures of neural networks in capturing human-like speech recognition
- AST: Audio Spectrogram Transformer
- Speech To Semantics: Improve ASR and NLU Jointly via All-Neural Interfaces
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound Localisation
- You Do Not Need More Data: Improving End-To-End Speech Recognition by Text-To-Speech Data Augmentation
- Few-Shot Keyword Spotting in Any Language
- Parameter Efficient Multimodal Transformers for Video Representation Learning
- End-to-end Neural Diarization: From Transformer to Conformer
- Deep speech inpainting of time-frequency masks
- Voice activity detection in the wild: A data-driven approach using teacher-student training
- Multimodal Fish Feeding Intensity Assessment in Aquaculture
- A Practical Deep Learning-Based Acoustic Side Channel Attack on Keyboards
- RNN-T For Latency Controlled ASR With Improved Beam Search
- Audio Captioning Transformer
- Augmentation Methods on Monophonic Audio for Instrument Classification in Polyphonic Music
- A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition
- Data Augmentation for End-to-end Code-switching Speech Recognition
- SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation
- QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions
- LEAF: A Learnable Frontend for Audio Classification
- Transformer Transducer: One Model Unifying Streaming and Non-streaming Speech Recognition
- Multi-Format Contrastive Learning of Audio Representations
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech Recognition
- Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR
- CLOCS: Contrastive Learning of Cardiac Signals Across Space, Time, and Patients
- BERTphone: Phonetically-Aware Encoder Representations for Utterance-Level Speaker and Language Recognition
- Non-autoregressive Transformer-based End-to-end ASR using BERT
- Tied & Reduced RNN-T Decoder
- Iterative Pseudo-Labeling for Speech Recognition
- The Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap
- N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
- Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
- TutorNet: Towards Flexible Knowledge Distillation for End-to-End Speech Recognition
- Improved Meta Learning for Low Resource Speech Recognition
- Semantic Mask for Transformer based End-to-End Speech Recognition
- TSNAT: Two-Step Non-Autoregressvie Transformer Models for Speech Recognition
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling
- A Four-Stage Data Augmentation Approach to ResNet-Conformer Based Acoustic Modeling for Sound Event Localization and Detection
- Imputer: Sequence Modelling via Imputation and Dynamic Programming
- Role of Data Augmentation Strategies in Knowledge Distillation for Wearable Sensor Data
- Layer-wise Fast Adaptation for End-to-End Multi-Accent Speech Recognition
- On Using SpecAugment for End-to-End Speech Translation
- The VoxCeleb Speaker Recognition Challenge: A Retrospective
- Analyzing analytical methods: The case of phonology in neural models of spoken language
- U2++: Unified Two-pass Bidirectional End-to-end Model for Speech Recognition
- NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
- Explaining Deep Classification of Time-Series Data with Learned Prototypes
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- EfficientTDNN: Efficient Architecture Search for Speaker Recognition
- Low Latency End-to-End Streaming Speech Recognition with a Scout Network
- Improving the Robustness and Clinical Applicability of Automatic Respiratory Sound Classification Using Deep Learning-Based Audio Enhancement: Algorithm Development and Validation
- RFBoost: Understanding and Boosting Deep WiFi Sensing via Physical Data Augmentation
- Audio Tagging by Cross Filtering Noisy Labels
- Unsupervised Speech Recognition
- Bayesian Learning for Deep Neural Network Adaptation
- The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
- NIST SRE CTS Superset: A large-scale dataset for telephony speaker recognition
- The Influence of Dataset Partitioning on Dysfluency Detection Systems
- An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning
- A Squeeze-and-Excitation and Transformer based Cross-task System for Environmental Sound Recognition
- Cross-Lingual Speaker Verification with Domain-Balanced Hard Prototype Mining and Language-Dependent Score Normalization
- Towards a Competitive End-to-End Speech Recognition for CHiME-6 Dinner Party Transcription
- Online Automatic Speech Recognition with Listen, Attend and Spell Model
- Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
- Conformer-based Target-Speaker Automatic Speech Recognition for Single-Channel Audio
- Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals
- Improving Speech Translation by Cross-Modal Multi-Grained Contrastive Learning
- Learning Representations for New Sound Classes With Continual Self-Supervised Learning
- On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
- Preech: A System for Privacy-Preserving Speech Transcription
- ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
- Multilingual Graphemic Hybrid ASR with Massive Data Augmentation
- Adaptive Weighting Scheme for Automatic Time-Series Data Augmentation
- Pushing the Limits of Non-Autoregressive Speech Recognition
- On Multitask Loss Function for Audio Event Detection and Localization
- MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition
- On Data-Augmentation and Consistency-Based Semi-Supervised Learning
- Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family Audio
- CL4AC: A Contrastive Loss for Audio Captioning
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
- Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation
- A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition
- CLAR: Contrastive Learning of Auditory Representations
- Data augmentation using prosody and false starts to recognize non-native children's speech
- The IDLAB VoxCeleb Speaker Recognition Challenge 2020 System Description
- W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
- DropDim: A Regularization Method for Transformer Networks
- Multimodal Urban Sound Tagging with Spatiotemporal Context
- Music theme recognition using CNN and self-attention
- SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- Exploring Wav2vec 2.0 fine-tuning for improved speech emotion recognition
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
- MAM: Masked Acoustic Modeling for End-to-End Speech-to-Text Translation
- MusiCoder: A Universal Music-Acoustic Encoder Based on Transformers
- Curriculum Pre-training for End-to-End Speech Translation
- Mel-spectrogram augmentation for sequence to sequence voice conversion
- Data Augmentation for Deep Learning-based Radio Modulation Classification
- Improving Mandarin End-to-End Speech Recognition with Word N-gram Language Model
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
- Learning Robust and Multilingual Speech Representations
- Multi-head Monotonic Chunkwise Attention For Online Speech Recognition
- Conditional Generative Data Augmentation for Clinical Audio Datasets
- Self-training and Pre-training are Complementary for Speech Recognition
- Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition
- Scaling End-to-End Models for Large-Scale Multilingual ASR
- MM-ALT: A Multimodal Automatic Lyric Transcription System
- Dynamic Acoustic Unit Augmentation With BPE-Dropout for Low-Resource End-to-End Speech Recognition
- Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition
- RNN-T Models Fail to Generalize to Out-of-Domain Audio: Causes and Solutions
- Heavily Augmented Sound Event Detection utilizing Weak Predictions
- ACCDOA: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization and Detection
- Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
- CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus
- Efficient conformer-based speech recognition with linear attention
- Listen with Intent: Improving Speech Recognition with Audio-to-Intent Front-End
- Dataset balancing can hurt model performance
- Multitask-Based Joint Learning Approach To Robust ASR For Radio Communication Speech
- SpecMix : A Mixed Sample Data Augmentation method for Training withTime-Frequency Domain Features
- Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information
- Predicting Rigid Body Dynamics using Dual Quaternion Recurrent Neural Networks with Quaternion Attention
- Sound Event Localization and Detection Using Activity-Coupled Cartesian DOA Vector and RD3net
- Data Augmentation For Children's Speech Recognition -- The "Ethiopian" System For The SLT 2021 Children Speech Recognition Challenge
- An investigation of phone-based subword units for end-to-end speech recognition
- Deep Residual Local Feature Learning for Speech Emotion Recognition
- Enforcing Encoder-Decoder Modularity in Sequence-to-Sequence Models
- Integrating Lattice-Free MMI into End-to-End Speech Recognition
- Universal ASR: Unifying Streaming and Non-Streaming ASR Using a Single Encoder-Decoder Model
- Exploring Continuous Integrate-and-Fire for Adaptive Simultaneous Speech Translation
- SCARF: Self-Supervised Contrastive Learning using Random Feature Corruption
- Differentiable Weighted Finite-State Transducers
- VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
- Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
- Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition
- Self-training with noisy student model and semi-supervised loss function for dcase 2021 challenge task 4
- Contrastive-mixup learning for improved speaker verification
- Supervised Contrastive Learning for Accented Speech Recognition
- Streaming End-to-End Bilingual ASR Systems with Joint Language Identification
- A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
- end-to-end training of a large vocabulary end-to-end speech recognition system
- SAN-M: Memory Equipped Self-Attention for End-to-End Speech Recognition
- WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition
- Bimodal Speech Emotion Recognition Using Pre-Trained Language Models
- Neural model robustness for skill routing in large-scale conversational AI systems: A design choice exploration
- Do End-to-End Speech Recognition Models Care About Context?
- Adversarial defense for automatic speaker verification by cascaded self-supervised learning models
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition
- The ByteDance Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2021
- Leveraging Phone Mask Training for Phonetic-Reduction-Robust E2E Uyghur Speech Recognition
- Transportation mode recognition based on low-rate acceleration and location signals with an attention-based multiple-instance learning network
- A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
- Unified End-to-End Speech Recognition and Endpointing for Fast and Efficient Speech Systems
- A Simple Baseline for Domain Adaptation in End to End ASR Systems Using Synthetic Data
- Training speaker recognition systems with limited data
- ASR-Aware End-to-end Neural Diarization
- NAS-VAD: Neural Architecture Search for Voice Activity Detection
- Swiss Parliaments Corpus, an Automatically Aligned Swiss German Speech to Standard German Text Corpus
- The xx205 System for the VoxCeleb Speaker Recognition Challenge 2020
- An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
- Attention based end to end Speech Recognition for Voice Search in Hindi and English
- Broadcasted Residual Learning for Efficient Keyword Spotting
- Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
- Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
- LiteDenseNet: A Lightweight Network for Hyperspectral Image Classification
- Leveraging End-to-End Speech Recognition with Neural Architecture Search
- Improved Conformer-based End-to-End Speech Recognition Using Neural Architecture Search
- Learning Speaker Embedding with Momentum Contrast
- End-to-End Spectro-Temporal Graph Attention Networks for Speaker Verification Anti-Spoofing and Speech Deepfake Detection
- Advanced Long-Content Speech Recognition With Factorized Neural Transducer
- XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition
- Mandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models
- Learnable Spectro-temporal Receptive Fields for Robust Voice Type Discrimination
- Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT
- CAPE: Encoding Relative Positions with Continuous Augmented Positional Embeddings
- Exploiting Spectral Augmentation for Code-Switched Spoken Language Identification
- Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech Translation
- Domain Generalization on Efficient Acoustic Scene Classification using Residual Normalization
- Improving accuracy of rare words for RNN-Transducer through unigram shallow fusion
- Intermediate Loss Regularization for CTC-based Speech Recognition
- A study of latent monotonic attention variants
- Keyword localisation in untranscribed speech using visually grounded speech models
- Cosine Scoring with Uncertainty for Neural Speaker Embedding
- Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
- Mitigating Unauthorized Speech Synthesis for Voice Protection
- Large-Scale Self- and Semi-Supervised Learning for Speech Translation
- Semi-supervised ASR by End-to-end Self-training
- Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
- Attention as a Guide for Simultaneous Speech Translation
- Class LM and word mapping for contextual biasing in End-to-End ASR
- HLT-NUS SUBMISSION FOR 2020 NIST Conversational Telephone Speech SRE
- TSUP Speaker Diarization System for Conversational Short-phrase Speaker Diarization Challenge
- Study of positional encoding approaches for Audio Spectrogram Transformers
- OLR 2021 Challenge: Datasets, Rules and Baselines
- BiQGEMM: Matrix Multiplication with Lookup Table For Binary-Coding-based Quantized DNNs
- EchoScan: Scanning Complex Room Geometries via Acoustic Echoes
- Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict
- SpeechNet: A Universal Modularized Model for Speech Processing Tasks
- FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization
- AutoKWS: Keyword Spotting with Differentiable Architecture Search
- Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information
- The NPU System for the 2020 Personalized Voice Trigger Challenge
- CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
- MixSpeech: Data Augmentation for Low-resource Automatic Speech Recognition
- Towards Semi-Supervised Semantics Understanding from Speech
- Improved Mask-CTC for Non-Autoregressive End-to-End ASR
- Robust Beam Search for Encoder-Decoder Attention Based Speech Recognition without Length Bias
- Domain Adaptation of low-resource Target-Domain models using well-trained ASR Conformer Models
- Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders
- Attention-based Transducer for Online Speech Recognition
- ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
- Detecting and analyzing missing citations to published scientific entities
- Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems
- On monoaural speech enhancement for automatic recognition of real noisy speech using mixture invariant training
- NeurST: Neural Speech Translation Toolkit
- Affective social anthropomorphic intelligent system
- An Investigation of the Effectiveness of Phase for Audio Classification
- Improving RNN transducer with normalized jointer network
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- ACGAN-based Data Augmentation Integrated with Long-term Scalogram for Acoustic Scene Classification
- Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification
- Distilling Knowledge from Ensembles of Acoustic Models for Joint CTC-Attention End-to-End Speech Recognition
- Deepfake Detection System for the ADD Challenge Track 3.2 Based on Score Fusion
- Multi-stage Progressive Compression of Conformer Transducer for On-device Speech Recognition
- Contextual Density Ratio for Language Model Biasing of Sequence to Sequence ASR Systems
- Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
- Spatial mixup: Directional loudness modification as data augmentation for sound event localization and detection
- Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation
- Romanian Speech Recognition Experiments from the ROBIN Project
- Unsupervised Domain Adaptation Schemes for Building ASR in Low-resource Languages
- Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition
- MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection
- Weak-Attention Suppression For Transformer Based Speech Recognition
- Logic-Guided Data Augmentation and Regularization for Consistent Question Answering
- Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers
- Word-level Embeddings for Cross-Task Transfer Learning in Speech Processing
- One In A Hundred: Select The Best Predicted Sequence from Numerous Candidates for Streaming Speech Recognition
- The HW-TSC's Offline Speech Translation Systems for IWSLT 2021 Evaluation
- Make More of Your Data: Minimal Effort Data Augmentation for Automatic Speech Recognition and Translation
- Attentional Speech Recognition Models Misbehave on Out-of-domain Utterances
- Dealing with training and test segmentation mismatch: FBK@IWSLT2021
- Beijing ZKJ-NPU Speaker Verification System for VoxCeleb Speaker Recognition Challenge 2021
- Improving End-To-End Modeling for Mispronunciation Detection with Effective Augmentation Mechanisms
- WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis
- Efficient Conformer with Prob-Sparse Attention Mechanism for End-to-EndSpeech Recognition
- Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
- Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and Conformers
- An evaluation of word-level confidence estimation for end-to-end automatic speech recognition
- Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
- Bayesian Learning of LF-MMI Trained Time Delay Neural Networks for Speech Recognition
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
- Boosting Active Learning for Speech Recognition with Noisy Pseudo-labeled Samples
- Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining
- EventDrop: data augmentation for event-based learning
- Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus
- Multi-Encoder-Decoder Transformer for Code-Switching Speech Recognition
- Momentum Pseudo-Labeling for Semi-Supervised Speech Recognition
- Distilling the Knowledge of BERT for Sequence-to-Sequence ASR
- CRNNs for Urban Sound Tagging with spatiotemporal context
- Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
- First Order Ambisonics Domain Spatial Augmentation for DNN-based Direction of Arrival Estimation
- The CUHK-TUDELFT System for The SLT 2021 Children Speech Recognition Challenge
- What shall we do with an hour of data? Speech recognition for the un- and under-served languages of Common Voice
- Towards Robust Waveform-Based Acoustic Models
- Decoupling Pronunciation and Language for End-to-end Code-switching Automatic Speech Recognition
- End-to-end Audio-visual Speech Recognition with Conformers
- Urban Sound Classification : striving towards a fair comparison
- Self-Supervised Representations Improve End-to-End Speech Translation
- OkwuGbé: End-to-End Speech Recognition for Fon and Igbo
- Decision Attentive Regularization to Improve Simultaneous Speech Translation Systems
- Data augmentation for learning predictive models on EEG: a systematic comparison
- Neural Model Reprogramming with Similarity Based Mapping for Low-Resource Spoken Command Recognition
- Convolutional Speech Recognition with Pitch and Voice Quality Features
- Shot Contrastive Self-Supervised Learning for Scene Boundary Detection
- Speech Emotion Recognition with Multiscale Area Attention and Data Augmentation
- Injecting Text in Self-Supervised Speech Pretraining
- Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation
- Contrastive Neural Processes for Self-Supervised Learning
- Data Techniques For Online End-to-end Speech Recognition
- SpecAugment on Large Scale Datasets
- Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
- Equivalence of Segmental and Neural Transducer Modeling: A Proof of Concept
- Orthros: Non-autoregressive End-to-end Speech Translation with Dual-decoder
- Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
- Investigation of Speaker-adaptation methods in Transformer based ASR
- Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone
- Learning with Out-of-Distribution Data for Audio Classification
- Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated Pseudotranscripts
- Streaming Multi-speaker ASR with RNN-T
- Robustness Testing of Language Understanding in Task-Oriented Dialog
- Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?
- Revisiting data augmentation for subspace clustering
- On a time-frequency blurring operator with applications in data augmentation
- TENET: A Time-reversal Enhancement Network for Noise-robust ASR
- The THUEE System Description for the IARPA OpenASR21 Challenge
- End-to-End Speaker Height and age estimation using Attention Mechanism with LSTM-RNN
- Speech Sentiment Analysis via Pre-trained Features from End-to-end ASR Models
- Low Resource German ASR with Untranscribed Data Spoken by Non-native Children -- INTERSPEECH 2021 Shared Task SPAPL System
- Insertion-Based Modeling for End-to-End Automatic Speech Recognition
- AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition
- Investigating the Reordering Capability in CTC-based Non-Autoregressive End-to-End Speech Translation
- Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
- Multistream CNN for Robust Acoustic Modeling
- ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition
- Reducing Streaming ASR Model Delay with Self Alignment
- Gated Recurrent Fusion with Joint Training Framework for Robust End-to-End Speech Recognition
- Residual Energy-Based Models for End-to-End Speech Recognition
- SpecRegMatch: Robust Semi-Supervised Regression for Vehicle Interior Noise Prediction
- Confidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition
- Semi-Supervised Learning with Data Augmentation for End-to-End ASR
- Investigation on Data Adaptation Techniques for Neural Named Entity Recognition
- End-to-end lyrics Recognition with Voice to Singing Style Transfer
- PDAugment: Data Augmentation by Pitch and Duration Adjustments for Automatic Lyrics Transcription
- Direct Models for Simultaneous Translation and Automatic Subtitling: FBK@IWSLT2023
- SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation
- power-law nonlinearity with maximally uniform distribution criterion for improved neural network training in automatic speech recognition
- KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition
- On Target Segmentation for Direct Speech Translation
- Howl: A Deployed, Open-Source Wake Word Detection System
- Mask Detection and Breath Monitoring from Speech: on Data Augmentation, Feature Representation and Modeling
- MASRI-HEADSET: A Maltese Corpus for Speech Recognition
- Capturing scattered discriminative information using a deep architecture in acoustic scene classification
- Polyphonic sound event detection based on convolutional recurrent neural networks with semi-supervised loss function for DCASE challenge 2020 task 4
- Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-based LVCSR
- Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
- Fusing information streams in end-to-end audio-visual speech recognition
- End-to-End Speaker-Attributed ASR with Transformer
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
- Contrastive Semi-supervised Learning for ASR
- End-to-End Automatic Speech Recognition with Deep Mutual Learning
- Content-Aware Speaker Embeddings for Speaker Diarisation
- Tiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devices
- EfficientNet-Absolute Zero for Continuous Speech Keyword Spotting
- REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling
- Self-supervised Text-independent Speaker Verification using Prototypical Momentum Contrastive Learning
- Data Augmentation with Locally-time Reversed Speech for Automatic Speech Recognition
- Back from the future: bidirectional CTC decoding using future information in speech recognition
- Significance of Data Augmentation for Improving Cleft Lip and Palate Speech Recognition
- Tree-constrained Pointer Generator for End-to-end Contextual Speech Recognition
- Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker
- A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English
- Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation Task
- The HCCL Speaker Verification System for Far-Field Speaker Verification Challenge
- Improving Polyphonic Sound Event Detection on Multichannel Recordings with the Sørensen-Dice Coefficient Loss and Transfer Learning
- E2E-based Multi-task Learning Approach to Joint Speech and Accent Recognition
- Impact of data-splits on generalization: Identifying COVID-19 from cough and context
- Efficient acoustic feature transformation in mismatched environments using a Guided-GAN
- Exploring Representation Learning for Small-Footprint Keyword Spotting
- Key Frame Mechanism For Efficient Conformer Based End-to-end Speech Recognition
- Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
- Towards Generating Diverse Audio Captions via Adversarial Training
- Improved Techniques for the Conditional Generative Augmentation of Clinical Audio Data
- Graph Structure Based Data Augmentation Method
- SpeechNAS: Towards Better Trade-off between Latency and Accuracy for Large-Scale Speaker Verification
- Deformable TDNN with adaptive receptive fields for speech recognition
- Scaling sparsemax based channel selection for speech recognition with ad-hoc microphone arrays
- Joint Representation Learning and Novel Category Discovery on Single- and Multi-modal Data
- Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
- The NTNU Taiwanese ASR System for Formosa Speech Recognition Challenge 2020
- Transformer-based end-to-end speech recognition with residual Gaussian-based self-attention
- A review of on-device fully neural end-to-end automatic speech recognition algorithms
- Deep Discriminative Feature Learning for Accent Recognition
- EasyASR: A Distributed Machine Learning Platform for End-to-end Automatic Speech Recognition
- Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning
- A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks
- A study on more realistic room simulation for far-field keyword spotting
- End-to-end Whispered Speech Recognition with Frequency-weighted Approaches and Pseudo Whisper Pre-training
- Cross-attention conformer for context modeling in speech enhancement for ASR
- Evaluating robustness of You Only Hear Once(YOHO) Algorithm on noisy audios in the VOICe Dataset
- DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes
- How Data Augmentation affects Optimization for Linear Regression
- Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer
- Effective Decoder Masking for Transformer Based End-to-End Speech Recognition
- Data Augmentation through Expert-guided Symmetry Detection to Improve Performance in Offline Reinforcement Learning
- AlignNet: A Unifying Approach to Audio-Visual Alignment
- Streaming Transformer ASR with Blockwise Synchronous Beam Search
- ESPnet-ST IWSLT 2021 Offline Speech Translation System
- Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications
- SapAugment: Learning A Sample Adaptive Policy for Data Augmentation
- Multi-style Training for South African Call Centre Audio
- On Comparison of Encoders for Attention based End to End Speech Recognition in Standalone and Rescoring Mode
- Improving RNN Transducer Based ASR with Auxiliary Tasks
- Two-stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding
- End-to-End Rich Transcription-Style Automatic Speech Recognition with Semi-Supervised Learning
- Synth2Aug: Cross-domain speaker recognition with TTS synthesized speech
- Acoustic span embeddings for multilingual query-by-example search
- Non-autoregressive Mandarin-English Code-switching Speech Recognition
- Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
- Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
- The USTC-NELSLIP Systems for Simultaneous Speech Translation Task at IWSLT 2021
- Simplified Self-Attention for Transformer-based End-to-End Speech Recognition
- Noisy Training Improves E2E ASR for the Edge
- SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding
- A Low-Compexity Deep Learning Framework For Acoustic Scene Classification
- Environment Transfer for Distributed Systems
- Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
- Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
- The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment
- Improving Streaming Automatic Speech Recognition With Non-Streaming Model Distillation On Unsupervised Data
- Integrating Source-channel and Attention-based Sequence-to-sequence Models for Speech Recognition
- Representation Learning for Sequence Data with Deep Autoencoding Predictive Components
- Improved Multi-Stage Training of Online Attention-based Encoder-Decoder Models
- Differentiable Allophone Graphs for Language-Universal Speech Recognition
- An Empirical Study of End-to-end Simultaneous Speech Translation Decoding Strategies
- What does a network layer hear? Analyzing hidden representations of end-to-end ASR through speech synthesis
- Attention-based ASR with Lightweight and Dynamic Convolutions
- From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation
- Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models
- MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
- Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
- Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers
- QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus
- Label-Synchronous Speech-to-Text Alignment for ASR Using Forward and Backward Transformers
- DeepSpectrumLite: A Power-Efficient Transfer Learning Framework for Embedded Speech and Audio Processing from Decentralised Data
- Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
- Similarity Analysis of Self-Supervised Speech Representations
- CAT: A CTC-CRF based ASR Toolkit Bridging the Hybrid and the End-to-end Approaches towards Data Efficiency and Low Latency
- Streaming Transformer-based Acoustic Models Using Self-attention with Augmented Memory
- On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models
- Self-Supervised Dynamic Networks for Covariate Shift Robustness
- LSTM and GPT-2 Synthetic Speech Transfer Learning for Speaker Recognition to Overcome Data Scarcity
- Memory-augmented conformer for improved end-to-end long-form ASR
- Efficient minimum word error rate training of RNN-Transducer for end-to-end speech recognition
- Should We Always Separate?: Switching Between Enhanced and Observed Signals for Overlapping Speech Recognition
- Embedded Emotions -- A Data Driven Approach to Learn Transferable Feature Representations from Raw Speech Input for Emotion Recognition
- Simultaneous Speech Translation for Live Subtitling: from Delay to Display
- A Visual Domain Transfer Learning Approach for Heartbeat Sound Classification
- Towards Data-efficient Modeling for Wake Word Spotting
- Searchable Hidden Intermediates for End-to-End Models of Decomposable Sequence Tasks
- Parallel Rescoring with Transformer for Streaming On-Device Speech Recognition
- Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions
- Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study
- CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription
- SNRi Target Training for Joint Speech Enhancement and Recognition
- Semi-supervised music emotion recognition using noisy student training and harmonic pitch class profiles
- Investigation of Training Label Error Impact on RNN-T
- Regularizing Recurrent Neural Networks via Sequence Mixup
- CIF-based Collaborative Decoding for End-to-end Contextual Speech Recognition
- Self-supervised Deep Learning for Reading Activity Classification
- Frame-level SpecAugment for Deep Convolutional Neural Networks in Hybrid ASR Systems
- Modeling Homophone Noise for Robust Neural Machine Translation
- Generalized Operating Procedure for Deep Learning: an Unconstrained Optimal Design Perspective
- Small energy masking for improved neural network training for end-to-end speech recognition
- Semi-supervised Sound Event Detection using Random Augmentation and Consistency Regularization
- Enhancing Audio Augmentation Methods with Consistency Learning
- Hierarchical Transformer-based Large-Context End-to-end ASR with Large-Context Knowledge Distillation
- On Compositions of Transformations in Contrastive Self-Supervised Learning
- Gaussian Kernelized Self-Attention for Long Sequence Data and Its Application to CTC-based Speech Recognition
- Unit selection synthesis based data augmentation for fixed phrase speaker verification
- Mutually-Constrained Monotonic Multihead Attention for Online ASR
- Streaming Multi-talker Speech Recognition with Joint Speaker Identification
- Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency
- FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip Regularization
- Learning Metrics from Mean Teacher: A Supervised Learning Method for Improving the Generalization of Speaker Verification System
- MCSAE: Masked Cross Self-Attentive Encoding for Speaker Embedding
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- Semantic Data Augmentation for End-to-End Mandarin Speech Recognition
- The NTNU System at the Interspeech 2020 Non-Native Children's Speech ASR Challenge
- DD-CNN: Depthwise Disout Convolutional Neural Network for Low-complexity Acoustic Scene Classification
- Audiovisual transfer learning for audio tagging and sound event detection
- FastCorrect 2: Fast Error Correction on Multiple Candidates for Automatic Speech Recognition
- Accent Recognition with Hybrid Phonetic Features
- Speech Summarization using Restricted Self-Attention
- Attention-Free Keyword Spotting
- Self-Attentive Multi-Layer Aggregation with Feature Recalibration and Normalization for End-to-End Speaker Verification System
- RCT: Random Consistency Training for Semi-supervised Sound Event Detection
- Linguistic Knowledge in Data Augmentation for Natural Language Processing: An Example on Chinese Question Matching
- A comparison of streaming models and data augmentation methods for robust speech recognition
- Alternate Intermediate Conditioning with Syllable-level and Character-level Targets for Japanese ASR
- Far-Field Automatic Speech Recognition
- Whole page recognition of historical handwriting
- Leveraging Cross-Utterance Context For ASR Decoding
- Mitigating Data Imbalance in Automated Speaking Assessment
- Generalized Spoofing Detection Inspired from Audio Generation Artifacts
- Extremely Low Footprint End-to-End ASR System for Smart Device
- Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
- HMM-Free Encoder Pre-Training for Streaming RNN Transducer
- Prototypical Contrastive Learning For Improved Few-Shot Audio Classification
- EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
- Chunkwise Aligners for Streaming Speech Recognition
- MSMT-FN: Multi-segment Multi-task Fusion Network for Marketing Audio Classification
- Weakly Supervised Phonological Features for Pathological Speech Analysis
- Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation
- Deep Learning-Based Acoustic Mosquito Detection in Noisy Conditions Using Trainable Kernels and Augmentations
- An empirical study of weakly supervised audio tagging embeddings for general audio representations
- Sequence-level self-learning with multiple hypotheses
- A practical two-stage training strategy for multi-stream end-to-end speech recognition
- Data Augmentation Methods for End-to-end Speech Recognition on Distant-Talk Scenarios
- SIGTYP 2021 Shared Task: Robust Spoken Language Identification
- RealTranS: End-to-End Simultaneous Speech Translation with Convolutional Weighted-Shrinking Transformer
- Layer Pruning on Demand with Intermediate CTC
- Multi-mode Transformer Transducer with Stochastic Future Context
- Enriching Under-Represented Named-Entities To Improve Speech Recognition Performance
- Multilingual Speech Translation with Unified Transformer: Huawei Noah's Ark Lab at IWSLT 2021
- Enrollment-less training for personalized voice activity detection
- The Volctrans Neural Speech Translation System for IWSLT 2021
- Raw Differentiable Architecture Search for Speech Deepfake and Spoofing Detection
- Tongji University Team for the VoxCeleb Speaker Recognition Challenge 2020
- Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems
- Population Based Training for Data Augmentation and Regularization in Speech Recognition
- Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech Recognition
- Unified Autoregressive Modeling for Joint End-to-End Multi-Talker Overlapped Speech Recognition and Speaker Attribute Estimation
- Oriental Language Recognition (OLR) 2020: Summary and Analysis
- Exploiting Single-Channel Speech For Multi-channel End-to-end Speech Recognition
- Improving Speech Recognition Accuracy of Local POI Using Geographical Models
- The NiuTrans End-to-End Speech Translation System for IWSLT 2021 Offline Task
- Multi-path Convolutional Neural Networks Efficiently Improve Feature Extraction in Continuous Adventitious Lung Sound Detection
- Conformer-based End-to-end Speech Recognition With Rotary Position Embedding
- CarneliNet: Neural Mixture Model for Automatic Speech Recognition
- An Improved Single Step Non-autoregressive Transformer for Automatic Speech Recognition
- Attention based on-device streaming speech recognition with large speech corpus
- Reducing Exposure Bias in Training Recurrent Neural Network Transducers
- CASS-NAT: CTC Alignment-based Single Step Non-autoregressive Transformer for Speech Recognition
- 4-bit Quantization of LSTM-based Speech Recognition Models
- Coarse-To-Fine And Cross-Lingual ASR Transfer
- Self-Attention Channel Combinator Frontend for End-to-End Multichannel Far-field Speech Recognition
- Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented Speech
- Is "moby dick" a Whale or a Bird? Named Entities and Terminology in Speech Translation
- Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph
- ChannelAugment: Improving generalization of multi-channel ASR by training with input channel randomization
- Topic Model Robustness to Automatic Speech Recognition Errors in Podcast Transcripts
- Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates
- ASR Rescoring and Confidence Estimation with ELECTRA
- Integrating Categorical Features in End-to-End ASR
- Duality Temporal-channel-frequency Attention Enhanced Speaker Representation Learning
- Learning Models for Query by Vocal Percussion: A Comparative Study
- Cross-lingual Transfer for Speech Processing using Acoustic Language Similarity
- STC speaker recognition systems for the NIST SRE 2021
- Whole-Word Segmental Speech Recognition with Acoustic Word Embeddings
- MaxVA: Fast Adaptation of Step Sizes by Maximizing Observed Variance of Gradients
- Peak Detection On Data Independent Acquisition Mass Spectrometry Data With Semisupervised Convolutional Transformers
- Multitask Learning and Joint Optimization for Transformer-RNN-Transducer Speech Recognition
- Focus on the present: a regularization method for the ASR source-target attention layer
- Benchmarking LF-MMI, CTC and RNN-T Criteria for Streaming ASR
- Cascade RNN-Transducer: Syllable Based Streaming On-device Mandarin Speech Recognition with a Syllable-to-Character Converter
- Exploring End-to-End Multi-channel ASR with Bias Information for Meeting Transcription