activity
20242026
collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech

Xuanjun Chen, Yun-Shing Wu, Wei-Chung Lu +4

Recent neural audio codec-based speech generation (CodecFake) produces highly realistic audio, posing a challenge to existing deepfake countermeasure models. While using codec resy…

cs.SD2026

CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech

Xuanjun Chen, Jiawei Du, Haibin Wu +8

With the rapid advancement of neural audio codecs, codec-based speech generation (CoSG) systems have become highly powerful. Unfortunately, CoSG also enables the creation of highly…

cs.SD2026

Training-Efficient Text-to-Music Generation with State-Space Modeling

Wei-Jaw Lee, Fang-Chih Hsieh, Xuanjun Chen +2

Recent advances in text-to-music generation (TTM) have yielded high-quality results, but often at the cost of extensive compute and the use of large proprietary internal data. To i…

cs.SD2026

How Does Instrumental Music Help SingFake Detection?

Xuanjun Chen, Chia-Yu Hu, I-Ming Lin +8

Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how inst…

cs.SD2025

Exploring State-Space-Model based Language Model in Music Generation

Wei-Jaw Lee, Fang-Chih Hsieh, Xuanjun Chen +2

The recent surge in State Space Models (SSMs), particularly the emergence of Mamba, has established them as strong alternatives or complementary modules to Transformers across dive…

cs.SD2024

DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset

Jiawei Du, I-Ming Lin, I-Hsiang Chiu +6

Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, hu…