activity
20202026
most citedLightweight End-to-end Text-to-speech Synthesis for low resource on-device applications

6 citations · 9 across the 8 of their papers we have counts for

collaborators

8 papers

eess.AS2026

Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis

Biel Tura Vecino, Yoach Lacombe, Julian Weber +4

Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpolating between conditioned and un…

eess.AS2025★ 1 cited

Universal Semantic Disentangled Privacy-preserving Speech Representation Learning

Biel Tura Vecino, Subhadeep Maji, Aravind Varier +11

The use of audio recordings of human speech to train LLMs poses privacy concerns due to these models' potential to generate outputs that closely resemble artifacts in the training…

eess.AS2025

Investigating self-supervised features for expressive, multilingual voice conversion

Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Grzegorz Beringer +7

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between d…

cs.SD2025★ 6 cited

Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications

Biel Tura Vecino, Adam Gabryś, Daniel Mątwicki +4

Recent works have shown that modelling raw waveform directly from text in an end-to-end (E2E) fashion produces more natural-sounding speech than traditional neural text-to-speech (…

eess.AS2024★ 1 cited

Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations

Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Iván Vallés-Pérez +6

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities…

eess.AS2023

Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech

Guangyan Zhang, Thomas Merritt, Manuel Sam Ribeiro +10

Neural text-to-speech systems are often optimized on L1/L2 losses, which make strong assumptions about the distributions of the target data space. Aiming to improve those assumptio…