87 citations · 305 across the 15 of their papers we have counts for
19 papers
Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts
Detai Xin, Sharath Adavanne, Federico Ang +3
We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context…
Differentiable Tracking-Based Training of Deep Learning Sound Source Localizers
Sharath Adavanne, Archontis Politis, Tuomas Virtanen
Data-based and learning-based sound source localization (SSL) has shown promising results in challenging conditions, and is commonly set as a classification or a regression problem…
A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection
Archontis Politis, Sharath Adavanne, Daniel Krause +3
This report presents the dataset and baseline of Task 3 of the DCASE2021 Challenge on Sound Event Localization and Detection (SELD). The dataset is based on emulation of real recor…
Non-native English lexicon creation for bilingual speech synthesis
Arun Baby, Pranav Jawale, Saranya Vinnaitherthan +3
Bilingual English speakers speak English as one of their languages. Their English is of a non-native kind, and their conversations are of a code-mixed fashion. The intelligibility…
Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019
Archontis Politis, Annamaria Mesaros, Sharath Adavanne +2
Sound event localization and detection is a novel area of research that emerged from the combined interest of analyzing the acoustic scene in terms of the spatial and temporal acti…
A Dataset of Reverberant Spatial Sound Scenes with Moving Sources for Sound Event Localization and Detection
Archontis Politis, Sharath Adavanne, Tuomas Virtanen
This report presents the dataset and the evaluation setup of the Sound Event Localization & Detection (SELD) task for the DCASE 2020 Challenge. The SELD task refers to the problem…