Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
Zhiyuan Zhao, Lijian Lin, Ye Zhu +3
We present the LEMAS-Dataset, which, to our knowledge, is currently the largest open-source multilingual speech corpus with word-level timestamps. Covering over 150,000 hours acros…
cs.SD2025
MelTok: 2D Tokenization for Single-Codebook Audio Compression
Jingyi Li, Zhiyuan Zhao, Zhisheng Zhang +6
Large Audio Language Models (LALMs) have emerged with strong performance across diverse audio understanding tasks and can be further enhanced by neural audio codecs. Transitioning…
cs.SD2025
STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
Tao Feng, Zhiyuan Zhao, Yifan Xie +4
We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that r…