most citedELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

2 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech

Haibin Wu, Xiaofei Wang, Sefik Emre Eskimez +8

People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) syste…

eess.AS2024

TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers

Yakun Song, Zhuo Chen, Xiaofei Wang +3

Neural codec language model (LM) has demonstrated strong capability in zero-shot text-to-speech (TTS) synthesis. However, the codec LM often suffers from limitations in inference s…

eess.AS2024

Total-Duration-Aware Duration Modeling for Text-to-Speech Systems

Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker +9

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting t…

stat.ME2024

Inference of treatment effect and its regional modifiers using restricted mean survival time in multi-regional clinical trials

Kaiyuan Hua, Hwanhee Hong, Xiaofei Wang

Multi-regional clinical trials (MRCTs) play an increasingly crucial role in global pharmaceutical development by expediting data gathering and regulatory approval across diverse pa…

cs.CL20242 cited

ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Yakun Song, Zhuo Chen, Xiaofei Wang +2

The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, exi…