4 citations · 6 across the 4 of their papers we have counts for
6 papers
SpeechLMScore: Evaluating speech generation using speech language model
Soumi Maiti, Yifan Peng, Takaaki Saeki +1
While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality…
End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings
Soumi Maiti, Hakan Erdogan, Kevin Wilson +3
We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling spe…
Generating Multilingual Voices Using Speaker Space Translation Based on Bilingual Speaker Data
Soumi Maiti, Erik Marchi, Alistair Conkie
We present progress towards bilingual Text-to-Speech which is able to transform a monolingual voice to speak a second language while preserving speaker voice quality. We demonstrat…
Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement
Soumi Maiti, Michael I Mandel
Traditional speech enhancement systems produce speech with compromised quality. Here we propose to use the high quality speech generation capability of neural vocoders for better q…
Parametric Resynthesis with neural vocoders
Soumi Maiti, Michael I Mandel
Noise suppression systems generally produce output speech with compromised quality. We propose to utilize the high quality speech generation capability of neural vocoders for noise…
Speech denoising by parametric resynthesis
Soumi Maiti, Michael I Mandel
This work proposes the use of clean speech vocoder parameters as the target for a neural network performing speech enhancement. These parameters have been designed for text-to-spee…