activity
20172022
most citedVoice Activity Detection: Merging Source and Filter-based Information

82 citations · 103 across the 9 of their papers we have counts for

collaborators

10 papers

eess.AS20221 cited

Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise

Tuomo Raitio, Petko Petkov, Jiangchuan Li +3

We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise. The method c…

cs.CL2021

Combining speakers of multiple languages to improve quality of neural voices

Javier Latorre, Charlotte Bailleul, Tuuli Morrill +2

In this work, we explore multiple architectures and training procedures for developing a multi-speaker and multi-lingual neural TTS system with the goals of a) improving the qualit…

eess.AS20201 cited

Evaluating the Intelligibility Benefits of Neural Speech Enrichment for Listeners with Normal Hearing and Hearing Impairment using the Greek Harvard Corpus

Muhammed PV Shifas, Anna Sfakianaki, Theognosia Chimona +1

In this work we evaluate a neural based speech intelligibility booster based on spectral shaping and dynamic range compression (SSDRC), referred to as WaveNet-based SSDRC (wSSDRC),…

cs.SD20201 cited

Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion

Dipjyoti Paul, Muhammed PV Shifas, Yannis Pantazis +1

The increased adoption of digital assistants makes text-to-speech (TTS) synthesis systems an indispensable feature of modern mobile devices. It is hence desirable to build a system…

eess.AS2020

Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions

Dipjyoti Paul, Yannis Pantazis, Yannis Stylianou

Recent advancements in deep learning led to human-level performance in single-speaker speech synthesis. However, there are still limitations in terms of speech quality when general…

eess.AS202011 cited

A non-causal FFTNet architecture for speech enhancement

Muhammed PV Shifas, Nagaraj Adiga, Vassilis Tsiaras +1

In this paper, we suggest a new parallel, non-causal and shallow waveform domain architecture for speech enhancement based on FFTNet, a neural network for generating high quality a…