5 papers
Deepfake Audio Detection Using Self-supervised Fusion Representations
Khalid Zaman, Qixuan Huang, Muhammad Uzair +1
This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the…
Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection
Khalid Zaman, Masashi Unoki
Human-imitated speech poses a greater challenge than AI-generated speech for both human listeners and automatic detection systems. Unlike AI-generated speech, which often contains…
Noise-Aware In-Context Learning for Hallucination Mitigation in ALLMs
Qixuan Huang, Khalid Zaman, Masashi Unoki
Auditory large language models (ALLMs) have demonstrated strong general capabilities in audio understanding and reasoning tasks. However, their reliability is still undermined by h…
EM2LDL: A Multilingual Speech Corpus for Mixed Emotion Recognition through Label Distribution Learning
Xingfeng Li, Xiaohan Shi, Junjie Li +4
This study introduces EM2LDL, a novel multilingual speech corpus designed to advance mixed emotion recognition through label distribution learning. Addressing the limitations of pr…
Modeling and Estimation of Vocal Tract and Glottal Source Parameters Using ARMAX-LF Model
Kai Lia, Masato Akagia, Yongwei Lib +1
Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) mod…