activity
20242026
collaborators

7 papers

cs.SD2026

OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech

Yong Ren, Jiangyan Yi, Jianhua Tao +5

Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a dir…

cs.SD2025

P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation

Yong Ren, Jiangyan Yi, Tao Wang +7

Neural speech generation (NSG) has rapidly advanced as a key component of artificial intelligence-generated content, enabling the generation of high-quality, highly realistic speec…

cs.SD2025

ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection

Hao Gu, Jiangyan Yi, Chenglong Wang +6

Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large lang…

cs.HC2025

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

Zheng Lian, Haoyu Chen, Lan Chen +9

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level, from naive discriminative tasks to complex emotion unders…

cs.SD2024

Region-Based Optimization in Continual Learning for Audio Deepfake Detection

Yujie Chen, Jiangyan Yi, Cunhang Fan +10

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although…

cs.SD2024

Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio

Xinrui Yan, Jiangyan Yi, Jianhua Tao +6

Open environment oriented open set model attribution of deepfake audio is an emerging research topic, aiming to identify the generation models of deepfake audio. Most previous work…