collaborators

8 papers

cs.CL2026

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

Li Zhou, Lutong Yu, You Lyu +6

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongsi…

cs.SD2026

Scaling Speech Tokenizers with Diffusion Autoencoders

Yuancheng Wang, Zhenyu Tang, Yun Wang +9

Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understandi…

eess.AS2025

Leveraging Language Information for Target Language Extraction

Mehmet Sinan Yıldırım, Ruijie Tao, Wupeng Wang +2

Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory sy…

eess.AS2025

Audio Deepfake Verification

Li Wang, Junyi Ao, Linyong Gan +3

With the rapid development of deepfake technology, simply making a binary judgment of true or false on audio is no longer sufficient to meet practical needs. Accurately determining…

astro-ph.GA2025

Euclid: Quick Data Release (Q1) -- A census of dwarf galaxies across a range of distances and environments

F. R. Marleau, R. Habas, D. Carollo +204

The Euclid Q1 fields were selected for calibration purposes in cosmology and are therefore relatively devoid of nearby galaxies. However, this is precisely what makes them interest…

cs.SD2025

Overview of the Amphion Toolkit (v0.2)

Jiaqi Li, Xueyao Zhang, Yuancheng Wang +9

Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a v…