16 papers
Luna-TTS Family Technical Report
Feng Yin, Shuai Shi, Junjie Zheng +19
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulat…
Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution
Zhenglong Liu, Wangyou Zhang, Chenda Li +1
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts…
TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
Qinzhe Hu, Chenda Li, Wangyou Zhang +3
Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment…
ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
Masao Someki, Alexander Polok, Carlos Carvalho +14
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
Changhao Cheng, Wei Wang, Wangyou Zhang +4
Continuous speech representations based on Variational Autoencoders (VAEs) have emerged as a promising alternative to traditional spectrogram or discrete token based features for s…
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
Bing Han, Chushu Zhou, Yifan Yang +4
Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity…