activity
20232026
most citedEzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

1 citations · 2 across the 18 of their papers we have counts for

collaborators
Showing eess.ASShow all

11 papers · 1 filter

eess.AS2026

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

Jiarui Hai, Karan Thakkar, Ke Chen +5

Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions. However, exi…

eess.AS2025

FlexSED: Towards Open-Vocabulary Sound Event Detection

Jiarui Hai, Helin Wang, Weizhe Guo +1

Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fund…

eess.AS2025

SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering

Jiarui Hai, Mounya Elhilali

Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up…

eess.AS2025

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

Helin Wang, Jiarui Hai, Dading Chong +11

Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to…

eess.AS2025

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Helin Wang, Jiarui Hai, Dongchao Yang +7

Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary aud…

eess.AS2024★ 1 cited

EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Jiarui Hai, Yong Xu, Hao Zhang +4

We introduce EzAudio, a text-to-audio (T2A) generation framework designed to produce high-quality, natural-sounding sound effects. Core designs include: (1) We propose EzAudio-DiT,…