8 papers
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
Li Zhou, Lutong Yu, You Lyu +6
Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongsi…
Scaling Speech Tokenizers with Diffusion Autoencoders
Yuancheng Wang, Zhenyu Tang, Yun Wang +9
Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understandi…
Leveraging Language Information for Target Language Extraction
Mehmet Sinan Yıldırım, Ruijie Tao, Wupeng Wang +2
Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory sy…
Audio Deepfake Verification
Li Wang, Junyi Ao, Linyong Gan +3
With the rapid development of deepfake technology, simply making a binary judgment of true or false on audio is no longer sufficient to meet practical needs. Accurately determining…
Euclid: Quick Data Release (Q1) -- A census of dwarf galaxies across a range of distances and environments
F. R. Marleau, R. Habas, D. Carollo +204
The Euclid Q1 fields were selected for calibration purposes in cosmology and are therefore relatively devoid of nearby galaxies. However, this is precisely what makes them interest…
Overview of the Amphion Toolkit (v0.2)
Jiaqi Li, Xueyao Zhang, Yuancheng Wang +9
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a v…