6 citations · 10 across the 23 of their papers we have counts for
10 papers · 1 filter
Learning Time-Graph Frequency Representation for Monaural Speech Enhancement
Tingting Wang, Tianrui Wang, Meng Ge +2
The Graph Fourier Transform (GFT) has recently demonstrated promising results in speech enhancement. However, existing GFT-based speech enhancement approaches often employ fixed gr…
Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
Hanyu Meng, Vidhyasaharan Sethu, Eliathamby Ambikairajah +2
In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fix…
SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
Jinyang Wu, Nana Hou, Zihan Pan +3
The rapid growth of the digital economy in South-East Asia (SEA) has amplified the risks of audio deepfakes, yet current datasets cover SEA languages only sparsely, leaving models…
Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
Hexin Liu, Haoyang Zhang, Qiquan Zhang +4
Code-switching automatic speech recognition (CS-ASR) presents unique challenges due to language confusion introduced by spontaneous intra-sentence switching and accent bias that bl…
Benchmarking Gaslighting Attacks Against Speech Large Language Models
Jinyang Wu, Bin Zhu, Xiandong Zou +3
As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input beco…
Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study
Qiquan Zhang, Moran Chen, Zeyang Song +3
Advanced long-context modeling backbone networks, such as Transformer, Conformer, and Mamba, have demonstrated state-of-the-art performance in speech enhancement. However, a system…