activity
20242026
most citedSpectrogram features for audio and speech analysis

7 citations · 8 across the 3 of their papers we have counts for

collaborators

5 papers

eess.AS20267 cited

Spectrogram features for audio and speech analysis

Ian McLoughlin, Lam Pham, Yan Song +7

Spectrogram-based representations have grown to dominate the feature space for deep learning audio analysis systems, and are often adopted for speech analysis also. Initially, the…

cs.SD2025

CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation

Crystal Min Hui Poon, Pai Chet Ng, Xiaoxiao Miao +4

Instruction-guided text-to-speech (TTS) research has reached a maturity level where excellent speech generation quality is possible on demand, yet two coupled biases persist in red…

cs.SD2025

An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition

Haoyu Song, Ian McLoughlin, Qing Gu +2

Existing speech emotion recognition (SER) methods commonly suffer from the lack of high-quality large-scale corpus, partly due to the complex, psychological nature of emotion which…

cs.SD2025

Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries

Pengfei Cai, Yan Song, Qing Gu +3

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts hav…

cs.SD20241 cited

MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection

Pengfei Cai, Yan Song, Kang Li +2

Sound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges. However, they still r…