papers

Publications (65)

cs.SD2024

A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge

Xiaopeng Wang, Yi Lu, Xin Qi +4

This paper presents the development of a speech synthesis system for the LIMMITS'24 Challenge, focusing primarily on Track 2. The objective of the challenge is to establish a multi…

cs.SD2023

UnifySpeech: A Unified Framework for Zero-shot Text-to-Speech and Voice Conversion

Haogeng Liu, Tao Wang, Ruibo Fu +3

Text-to-speech (TTS) and voice conversion (VC) are two different tasks both aiming at generating high quality speaking voice according to different input modality. Due to their sim…

cs.SD2022

An Initial Investigation for Detecting Vocoder Fingerprints of Fake Audio

Xinrui Yan, Jiangyan Yi, Jianhua Tao +5

Many effective attempts have been made for fake audio detection. However, they can only provide detection results but no countermeasures to curb this harm. For many related practic…

cs.SD2024

Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy

Yuankun Xie, Ruibo Fu, Zhengqi Wen +5

With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID)…

cs.MM2025

Deconfounded Reasoning for Multimodal Fake News Detection via Causal Intervention

Moyang Liu, Kaiying Yan, Yukun Liu +4

The rapid growth of social media has led to the widespread dissemination of fake news across multiple content forms, including text, images, audio, and video. Traditional unimodal…

cs.SD2024

EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech

Xin Qi, Ruibo Fu, Zhengqi Wen +10

In the current era of Artificial Intelligence Generated Content (AIGC), a Low-Rank Adaptation (LoRA) method has emerged. It uses a plugin-based approach to learn new knowledge with…

cs.SD2024

Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0

Zhiyong Wang, Ruibo Fu, Zhengqi Wen +10

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and in…

cs.SD2024

A Noval Feature via Color Quantisation for Fake Audio Detection

Zhiyong Wang, Xiaopeng Wang, Yuankun Xie +9

In the field of deepfake detection, previous studies focus on using reconstruction or mask and prediction methods to train pre-trained models, which are then transferred to fake au…

cs.SD2023

TO-Rawnet: Improving RawNet with TCN and Orthogonal Regularization for Fake Audio Detection

Chenglong Wang, Jiangyan Yi, Jianhua Tao +4

Current fake audio detection relies on hand-crafted features, which lose information during extraction. To overcome this, recent studies use direct feature extraction from raw audi…

cs.SD2025

When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks

Zeyu Xie, Chenxing Li, Xuenan Xu +6

This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus em…

cs.SD2023

Half-Truth: A Partially Fake Audio Detection Dataset

Jiangyan Yi, Ye Bai, Jianhua Tao +5

Diverse promising datasets have been designed to hold back the development of fake audio detection, such as ASVspoof databases. However, previous datasets ignore an attacking situa…

cs.SD2022

Fully Automated End-to-End Fake Audio Detection

Chenglong Wang, Jiangyan Yi, Jianhua Tao +6

The existing fake audio detection systems often rely on expert experience to design the acoustic features or manually design the hyperparameters of the network structure. However,…

eess.AS2025

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

Bingsong Bai, Qihang Lu, Wenbing Yang +8

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets,…

cs.SD2024

Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge

Yuankun Xie, Xiaopeng Wang, Zhiyong Wang +4

ASVspoof5, the fifth edition of the ASVspoof series, is one of the largest global audio security challenges. It aims to advance the development of countermeasure (CM) to discrimina…

eess.AS2024

ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

Ruibo Fu, Xin Qi, Zhengqi Wen +10

Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media…

cs.SD2025

Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform

Yuankun Xie, Ruibo Fu, Xiaopeng Wang +5

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures…

cs.SD2024

The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio

Yuankun Xie, Yi Lu, Ruibo Fu +9

With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for generalized detection methods. ALM-based deepfake audio currently exhibits wi…

cs.SD2022

Singing-Tacotron: Global duration control attention and dynamic filter for End-to-end singing voice synthesis

Tao Wang, Ruibo Fu, Jiangyan Yi +2

End-to-end singing voice synthesis (SVS) is attractive due to the avoidance of pre-aligned data. However, the auto learned alignment of singing voice with lyrics is difficult to ma…

cs.MM2025

PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis

Heng Xie, Kang Zhu, Zhengqi Wen +4

Multimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating…

cs.MM2024

Exploring the Role of Audio in Multimodal Misinformation Detection

Moyang Liu, Yukun Liu, Ruibo Fu +4

With the rapid development of deepfake technology, especially the deep audio fake technology, misinformation detection on the social media scene meets a great challenge. Social med…

cs.SD2025

M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis

Xiaopeng Wang, Chunyu Qiang, Ruibo Fu +10

Non-autoregressive (NAR) text-to-speech synthesis relies on length alignment between text sequences and audio representations, constraining naturalness and expressiveness. Existing…

cs.SD2025

P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation

Yong Ren, Jiangyan Yi, Tao Wang +7

Neural speech generation (NSG) has rapidly advanced as a key component of artificial intelligence-generated content, enabling the generation of high-quality, highly realistic speec…

cs.SD2023

Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding

Chunyu Qiang, Hao Li, Hao Ni +5

Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations a…

cs.SD2024

ADD 2022: the First Audio Deep Synthesis Detection Challenge

Jiangyan Yi, Ruibo Fu, Jianhua Tao +17

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios.…

cs.CV2026

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation

Haojie Zhang, Zhihao Liang, Ruibo Fu +5

Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and computational efficiency. As video le…

cs.SD2024

Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection

Xiaopeng Wang, Ruibo Fu, Zhengqi Wen +9

The generalization of Fake Audio Detection (FAD) is critical due to the emergence of new spoofing techniques. Traditional FAD methods often focus solely on distinguishing between g…

cs.SD2023

CFAD: A Chinese Dataset for Fake Audio Detection

Haoxin Ma, Jiangyan Yi, Chenglong Wang +5

Fake audio detection is a growing concern and some relevant datasets have been designed for research. However, there is no standard public Chinese dataset under complex conditions.…

cs.SD2023

Adaptive Fake Audio Detection with Low-Rank Model Squeezing

Xiaohui Zhang, Jiangyan Yi, Jianhua Tao +3

The rapid advancement of spoofing algorithms necessitates the development of robust detection methods capable of accurately identifying emerging fake audio. Traditional approaches,…

eess.AS2024

Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation

Chenxu Xiong, Ruibo Fu, Shuchen Shi +9

Current mainstream audio generation methods primarily rely on simple text prompts, often failing to capture the nuanced details necessary for multi-style audio generation. To addre…

cs.SD2023

ADD 2023: the Second Audio Deepfake Detection Challenge

Jiangyan Yi, Jianhua Tao, Ruibo Fu +15

Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the…

eess.AS2025

SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec

Chunyu Qiang, Haoyu Wang, Cheng Gong +10

Speech codecs serve as a crucial bridge in unifying speech and text language models. Existing codec methods face several challenges in semantic encoding, such as residual paralingu…

cs.SD2024

Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?

Yuankun Xie, Chenxu Xiong, Xiaopeng Wang +9

Currently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs. These ALMs have significantly lowered the ba…

eess.AS2025

VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing

Chunyu Qiang, Wang Geng, Yi Zhao +12

Deep learning has brought significant improvements to the field of cross-modal representation learning. For tasks such as text-to-speech (TTS), voice conversion (VC), and automatic…

cs.CV2025

MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics

Cong Cai, Shan Liang, Xuefei Liu +11

Deception detection has garnered increasing attention in recent years due to the significant growth of digital media and heightened ethical and security concerns. It has been exten…

cs.LG2025

MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection

Kaiying Yan, Moyang Liu, Yukun Liu +5

Multimodal fake news detection is essential for maintaining the authenticity of Internet multimedia information. Significant differences in form and content of multimodal informati…

cs.SD2025

Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition

Yuankun Xie, Xiaopeng Wang, Zhiyong Wang +7

Current research in audio deepfake detection is gradually transitioning from binary classification to multi-class tasks, referred as audio deepfake source tracing task. However, ex…

cs.MM2025

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Le Xu, Chenxing Li, Yong Ren +5

Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge t…

cs.SD2024

Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation

Hongming Guo, Ruibo Fu, Yizhong Geng +9

Text-to-audio (TTA) model is capable of generating diverse audio from textual prompts. However, most mainstream TTA models, which predominantly rely on Mel-spectrograms, still face…

cs.SD2025

RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection

Ruibo Fu, Xiaopeng Wang, Zhengqi Wen +8

Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attac…

eess.AS2026

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

Chunyu Qiang, Xiaopeng Wang, Kang Yin +11

Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous…

cs.SD2024

Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection

Cunhang Fan, Mingming Ding, Jianhua Tao +4

Most research in synthetic speech detection (SSD) focuses on improving performance on standard noise-free datasets. However, in actual situations, noise interference is usually pre…

cs.CL2024

Fake News Detection and Manipulation Reasoning via Large Vision-Language Models

Ruihan Jin, Ruibo Fu, Zhengqi Wen +3

Fake news becomes a growing threat to information security and public opinion with the rapid sprawl of media manipulation. Therefore, fake news detection attracts widespread attent…

cs.SD2024

DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

Xin Qi, Ruibo Fu, Zhengqi Wen +12

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have…

cs.SD2026

DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer

Yisu Liu, Chenxing Li, Wanqian Zhang +6

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and…

cs.SD2026

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

Yuankun Xie, Haonan Cheng, Jiayi Zhou +10

The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including so…

cs.CL2025

Towards Diverse and Efficient Audio Captioning via Diffusion Models

Manjie Xu, Chenxing Li, Xinyi Tu +4

We introduce Diffusion-based Audio Captioning (DAC), a non-autoregressive diffusion model tailored for diverse and efficient audio captioning. Although existing captioning models r…

cs.MM2025

Exploring Modality Disruption in Multimodal Fake News Detection

Moyang Liu, Kaiying Yan, Yukun Liu +4

The rapid growth of social media has led to the widespread dissemination of fake news across multiple content forms, including text, images, audio, and video. Compared to unimodal…

cs.SD2026

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

Yuankun Xie, Ruibo Fu, Zhiyong Wang +5

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia se…

cs.SD2024

SceneFake: An Initial Dataset and Benchmarks for Scene Fake Audio Detection

Jiangyan Yi, Chenglong Wang, Jianhua Tao +5

Many datasets have been designed to further the development of fake audio detection. However, fake utterances in previous datasets are mostly generated by altering timbre, prosody,…

eess.AS2024

MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation

Ruibo Fu, Shuchen Shi, Hongming Guo +12

Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements…

cs.SD2024

Generalized Fake Audio Detection via Deep Stable Learning

Zhiyong Wang, Ruibo Fu, Zhengqi Wen +9

Although current fake audio detection approaches have achieved remarkable success on specific datasets, they often fail when evaluated with datasets from different distributions. P…

cs.SD2022

Emotion Selectable End-to-End Text-based Speech Editing

Tao Wang, Jiangyan Yi, Ruibo Fu +3

Text-based speech editing allows users to edit speech by intuitively cutting, copying, and pasting text to speed up the process of editing speech. In the previous work, CampNet (co…

eess.AS2023

Learning Speech Representation From Contrastive Token-Acoustic Pretraining

Chunyu Qiang, Hao Li, Yixin Tian +4

For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate…

cs.SD2022

An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era

Andreas Triantafyllopoulos, Björn W. Schuller, Gökçe İymen +9

Speech is the fundamental mode of human communication, and its synthesis has long been a core priority in human-computer interaction research. In recent years, machines have manage…

cs.SD2022

Text Enhancement for Paragraph Processing in End-to-End Code-switching TTS

Chunyu Qiang, Jianhua Tao, Ruibo Fu +4

Current end-to-end code-switching Text-to-Speech (TTS) can already generate high quality two languages speech in the same utterance with single speaker bilingual corpora. When the…

eess.AS2025

InstructAudio: Unified speech and music generation with natural language instruction

Chunyu Qiang, Kang Yin, Xiaopeng Wang +8

Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only…

eess.AS2024

ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024

Ruibo Fu, Rui Liu, Chunyu Qiang +11

The Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC 2024) is part of the ISCSLP 2024 Competitions and Challenges track. While current text-to-speech (TTS) techn…

cs.CL2025

Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning

Kaiying Yan, Moyang Liu, Yukun Liu +4

The rapid spread of fake news across multimedia platforms presents serious challenges to information credibility. In this paper, we propose a Debunk-and-Infer framework for Fake Ne…

cs.SD2022

CampNet: Context-Aware Mask Prediction for End-to-End Text-Based Speech Editing

Tao Wang, Jiangyan Yi, Ruibo Fu +2

The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major d…

cs.SD2022

NeuralDPS: Neural Deterministic Plus Stochastic Model with Multiband Excitation for Noise-Controllable Waveform Generation

Tao Wang, Ruibo Fu, Jiangyan Yi +2

The traditional vocoders have the advantages of high synthesis efficiency, strong interpretability, and speech editability, while the neural vocoders have the advantage of high syn…

cs.SD2024

PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation

Shuchen Shi, Ruibo Fu, Zhengqi Wen +10

Text-to-Audio (TTA) aims to generate audio that corresponds to the given text description, playing a crucial role in media production. The text descriptions in TTA datasets lack ri…

cs.MM2025

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Yong Ren, Chenxing Li, Le Xu +7

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remain…

cs.SD2023

Low-rank Adaptation Method for Wav2vec2-based Fake Audio Detection

Chenglong Wang, Jiangyan Yi, Xiaohui Zhang +3

Self-supervised speech models are a rapidly developing research topic in fake audio detection. Many pre-trained models can serve as feature extractors, learning richer and higher-l…

cs.SD2024

Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio

Yi Lu, Yuankun Xie, Ruibo Fu +9

With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typic…

cs.SD2026

Interpretable All-Type Audio Deepfake Detection with Audio LLMs via Frequency-Time Reinforcement Learning

Yuankun Xie, Xiaoxuan Guo, Jiayi Zhou +6

Recent advances in audio large language models (ALLMs) have made high-quality synthetic audio widely accessible, increasing the risk of malicious audio deepfakes across speech, env…