activity
20242026
collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

Yuxiang Wang, Hongyu Liu, Yijiang Xu +9

As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speak…

cs.SD2025

DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components

Yupei Li, Li Wang, Yuxiang Wang +5

Audio deepfake detection has recently garnered public concern due to its implications for security and reliability. Traditional deep learning methods have been widely applied to th…

cs.SD2025

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

Xueyao Zhang, Chaoren Wang, Huan Liao +8

Aligning large generative models with human feedback is a critical challenge. In speech synthesis, this is particularly pronounced due to the lack of a large-scale human preference…

cs.SD2025

Overview of the Amphion Toolkit (v0.2)

Jiaqi Li, Xueyao Zhang, Yuancheng Wang +9

Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a v…

cs.SD2024

Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning

Haorui He, Yuchen Song, Yuancheng Wang +6

The effectiveness of one-shot voice conversion (VC) decreases in real-world scenarios where reference speeches, which are often sourced from the internet, contain various disturban…