3 papers
cs.SD2025
When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection
Bin Hu, Kunyang Huang, Daehan Kwak +2
The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. W…
cs.SD2025
Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech
Kunyang Huang, Bin Hu
The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications…
cs.SD2024
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
Fan Zhang, Siyuan Zhao, Naye Ji +12
Speech-driven gesture generation using transformer-based generative models represents a rapidly advancing area within virtual human creation. However, existing models face signific…