Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge Evaluation Plan
Xueping Zhang, Han Yin, Yang Xiao +4
Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion…
cs.SD2026
MoEScore: Mixture-of-Experts-Based Text-Audio Relevance Score Prediction for Text-to-Audio System Evaluation
Bochao Sun, Yang Xiao, Han Yin
Recent advances in generative models have enabled modern Text-to-Audio (TTA) systems to synthesize audio with high perceptual quality. However, TTA systems often struggle to mainta…
cs.SD2025
Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
Han Yin, Jung-Woo Choi
Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To eval…