3 papers
cs.CL2026
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
cs.SD2026
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
Tao Li, Wenshuo Ge, Zhichao Wang +6
Codec-based language models (LMs) have revolutionized text-to-speech (TTS). However, standard codecs entangle timbre and prosody, which hinders independent control in continuation-…
cs.LG2025
HESSO: Towards Automatic Efficient and User Friendly Any Neural Network Training and Pruning
Tianyi Chen, Xiaoyi Qu, David Aponte +7
Structured pruning is one of the most popular approaches to effectively compress the heavy deep neural networks (DNNs) into compact sub-networks while retaining performance. The ex…