collaborators

8 papers

cs.CL2026

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +398

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…

cs.SD2026

Frequency-Aware Self-Supervised Music Representation Learning

Yicheng Gu, Junan Zhang, Jerry Li +2

Self-supervised learning (SSL) has emerged as an essential paradigm for music information retrieval (MIR). While current SSL models achieve state-of-the-art performance across vari…

cs.SD2026

Aliasing-Free Neural Audio Synthesis

Yicheng Gu, Junan Zhang, Chaoren Wang +3

In neural audio synthesis, neural vocoders and codecs are models that reconstruct waveforms from acoustic and latent representations, which are essential to the resulting audio qua…

eess.AS2026

Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech

Zirui Li, Jens Edlund, Yicheng Gu +3

Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. We present Nord-Pa…

cs.SD2025

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

Haorui He, Zengqiang Shang, Chaoren Wang +11

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in…

cs.SD2025

Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN

Yicheng Gu, Chaoren Wang, Zhizheng Wu +1

Pitch manipulation is the process of producers adjusting the pitch of an audio segment to a specific key and intonation, which is essential in music production. Neural-network-base…