collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2026

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

Yunjia Li, Menglin Wu, Junyu Dai +13

Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are tra…

eess.AS2026

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

Bajian Xiang, Cheng Wen, Han Zhao +12

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, au…

eess.AS2026

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

Haoxu Wang, Biao Tian, Weiqin Li +3

Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO,…

eess.AS2026

LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization

Zexu Pan, Shengkui Zhao, Yukun Ma +4

Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural au…

eess.AS2026

MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis

Keyu An, Zhiyu Zhang, Changfeng Gao +7

This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram fram…

eess.AS2026

FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning

Haoxu Wang, Biao Tian, Yiheng Jiang +5

Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-tr…