activity
20222026
most citedAn Initial Investigation for Detecting Vocoder Fingerprints of Fake Audio

28 citations · 57 across the 30 of their papers we have counts for

collaborators
Showing 2024 · cs.SDShow all

5 papers · 2 filters

cs.SD2024

Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation

Hongming Guo, Ruibo Fu, Yizhong Geng +9

Text-to-audio (TTA) model is capable of generating diverse audio from textual prompts. However, most mainstream TTA models, which predominantly rely on Mel-spectrograms, still face…

cs.SD2024

DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

Xin Qi, Ruibo Fu, Zhengqi Wen +12

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have…

cs.SD2024

WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification

Junzuo Zhou, Jiangyan Yi, Yong Ren +3

Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before co…

cs.SD2024

PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation

Shuchen Shi, Ruibo Fu, Zhengqi Wen +10

Text-to-Audio (TTA) aims to generate audio that corresponds to the given text description, playing a crucial role in media production. The text descriptions in TTA datasets lack ri…

cs.SD2024

TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking

Junzuo Zhou, Jiangyan Yi, Tao Wang +5

Various threats posed by the progress in text-to-speech (TTS) have prompted the need to reliably trace synthesized speech. However, contemporary approaches to this task involve add…