18 papers
Luna-TTS Family Technical Report
Feng Yin, Shuai Shi, Junjie Zheng +19
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulat…
Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution
Zhenglong Liu, Wangyou Zhang, Chenda Li +1
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts…
TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
Qinzhe Hu, Chenda Li, Wangyou Zhang +3
Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment…
ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
Masao Someki, Alexander Polok, Carlos Carvalho +14
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
Bing Han, Chushu Zhou, Yifan Yang +4
Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity…
SLM-SS: Speech Language Model for Generative Speech Separation
Tianhua Li, Chenda Li, Wei Wang +4
Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to…