activity
20172022
most citedLarge Batch Training of Convolutional Networks

507 citations · 761 across the 20 of their papers we have counts for

collaborators

30 papers

eess.AS2022

Adapter-Based Extension of Multi-Speaker Text-to-Speech Model for New Speakers

Cheng-Ping Hsieh, Subhankar Ghosh, Boris Ginsburg

Fine-tuning is a popular method for adapting text-to-speech (TTS) models to new speakers. However this approach has some challenges. Usually fine-tuning requires several hours of h…

cs.SD2022

Damage Control During Domain Adaptation for Transducer Based Automatic Speech Recognition

Somshubra Majumdar, Shantanu Acharya, Vitaly Lavrukhin +1

Automatic speech recognition models are often adapted to improve their accuracy in a new domain. A potential drawback of model adaptation to new domains is catastrophic forgetting,…

eess.AS20224 cited

Multi-scale Speaker Diarization with Dynamic Scale Weighting

Tae Jin Park, Nithin Rao Koluguri, Jagadeesh Balam +1

Speaker diarization systems are challenged by a trade-off between the temporal resolution and the fidelity of the speaker representation. By obtaining a superior temporal resolutio…

cs.CL2022

Shallow Fusion of Weighted Finite-State Transducer and Language Model for Text Normalization

Evelina Bakhturina, Yang Zhang, Boris Ginsburg

Text normalization (TN) systems in production are largely rule-based using weighted finite-state transducers (WFST). However, WFST-based systems struggle with ambiguous input when…

eess.AS2021

Mixer-TTS: non-autoregressive, fast and compact text-to-speech model conditioned on language model embeddings

Oktai Tatanov, Stanislav Beliaev, Boris Ginsburg

This paper describes Mixer-TTS, a non-autoregressive model for mel-spectrogram generation. The model is based on the MLP-Mixer architecture adapted for speech synthesis. The basic…

eess.AS20214 cited

TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Nithin Rao Koluguri, Taejin Park, Boris Ginsburg

In this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excit…