activity
20182026
most citedTowards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2026

SimpleTool: Parallel Decoding for Real-Time LLM Function Calling

Xiaoxin Shi, Jiaxin Wan, Linkang Dong +3

LLM-based function calling enables intelligent agents to interact with external tools and environments, yet autoregressive decoding imposes a fundamental latency bottleneck that li…

cs.SD2025

Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling

Junjie Cao, Yichen Han, Ruonan Zhang +5

Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation…

cs.SD2025

MBCodec:Thorough disentangle for high-fidelity audio compression

Ruonan Zhang, Xiaoyang Hao, Yichen Han +3

High-fidelity neural audio codecs in Text-to-speech (TTS) aim to compress speech signals into discrete representations for faithful reconstruction. However, prior approaches faced…

cs.SD2025

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

Yichen Han, Xiaoyang Hao, Keming Chen +25

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, whi…

cs.SD20251 cited

Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement

Qianniu Chen, Xiaoyang Hao, Bowen Li +2

Zero-shot Text-To-Speech (TTS) synthesis shows great promise for personalized voice customization through voice cloning. However, current methods for achieving zero-shot TTS heavil…

cs.LG2018

An initial attempt of combining visual selective attention with deep reinforcement learning

Liu Yuezhang, Ruohan Zhang, Dana H. Ballard

Visual attention serves as a means of feature selection mechanism in the perceptual system. Motivated by Broadbent's leaky filter model of selective attention, we evaluate how such…