collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2025

Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget

Andy T. Liu, Yi-Cheng Lin, Haibin Wu +2

Despite their impressive success, training foundation models remains computationally costly. This paper investigates how to efficiently train speech foundation models with self-sup…

eess.AS2024

EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations

Wenze Ren, Yi-Cheng Lin, Huang-Cheng Chou +5

The neural codec model reduces speech data transmission delay and serves as the foundational tokenizer for speech language models (speech LMs). Preserving emotional information in…

eess.AS2024

CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems

Haibin Wu, Yuan Tseng, Hung-yi Lee

Current state-of-the-art (SOTA) codec-based audio synthesis systems can mimic anyone's voice with just a 3-second sample from that specific unseen speaker. Unfortunately, malicious…

eess.AS2024

Singing Voice Graph Modeling for SingFake Detection

Xuanjun Chen, Haibin Wu, Jyh-Shing Roger Jang +1

Detecting singing voice deepfakes, or SingFake, involves determining the authenticity and copyright of a singing voice. Existing models for speech deepfake detection have struggled…

eess.AS2024

Neural Codec-based Adversarial Sample Detection for Speaker Verification

Xuanjun Chen, Jiawei Du, Haibin Wu +2

Automatic Speaker Verification (ASV), increasingly used in security-critical applications, faces vulnerabilities from rising adversarial attacks, with few effective defenses availa…