collaborators

9 papers

eess.AS2026

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

Yunjia Li, Menglin Wu, Junyu Dai +13

Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are tra…

eess.AS2026

Qwen-Audio-3.0-Gen-Preview Technical Report

Junyu Dai, Xiaoyue Duan, Xinyue Fan +14

The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…

eess.AS2026

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

Bajian Xiang, Cheng Wen, Han Zhao +12

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, au…

cs.SD2026

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

Junyu Dai, Xinyue Fan, Weiqin Li +14

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The propo…

eess.AS2026

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

Haoxu Wang, Biao Tian, Weiqin Li +3

Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO,…

cs.SD2026

Eliminating stability hallucinations in llm-based tts models via attention guidance

ShiMing Wang, ZhiHao Du, Yang Xiang +6

This paper focuses on resolving stability hallucinations (e.g., repetitive or omitted speech) in LLM-based Text-to-Speech (TTS) models by improving and leveraging the attention mec…