activity
20182026
most citedImproving Readability for Automatic Speech Recognition Transcription

19 citations · 34 across the 23 of their papers we have counts for

collaborators
Showing eess.ASShow all

17 papers · 1 filter

eess.AS2026

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

Freeman Jiang, Ramon Sanabria, Soham Deshmukh +16

Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited du…

eess.AS2024

TS3-Codec: Transformer-Based Simple Streaming Single Codec

Haibin Wu, Naoyuki Kanda, Sefik Emre Eskimez +1

Neural audio codecs (NACs) have garnered significant attention as key technologies for audio compression as well as audio representation for speech language models. While mainstrea…

eess.AS2024

Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech

Haibin Wu, Xiaofei Wang, Sefik Emre Eskimez +8

People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) syste…

eess.AS2024

An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS

Xiaofei Wang, Sefik Emre Eskimez, Manthan Thakker +8

Recently, zero-shot text-to-speech (TTS) systems, capable of synthesizing any speaker's voice from a short audio prompt, have made rapid advancements. However, the quality of the g…

eess.AS2024

Total-Duration-Aware Duration Modeling for Text-to-Speech Systems

Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker +9

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting t…

eess.AS2024

E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker +10

This paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-a…