most citedSparks of Large Audio Models: A Survey and Outlook

11 citations · 28 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SD202311 cited

Sparks of Large Audio Models: A Survey and Outlook

Siddique Latif, Moazzam Shoukat, Fahad Shamshad +8

This survey paper provides a comprehensive overview of the recent advancements and challenges in applying large language models to the field of audio signal processing. Audio proce…

cs.CL20238 cited

Cross-Language Speech Emotion Recognition Using Multimodal Dual Attention Transformers

Syed Aun Muhammad Zaidi, Siddique Latif, Junaid Qadir

Despite the recent progress in speech emotion recognition (SER), state-of-the-art systems are unable to achieve improved performance in cross-language settings. In this paper, we p…

cs.SD2023

A Preliminary Study on Augmenting Speech Emotion Recognition using a Diffusion Model

Ibrahim Malik, Siddique Latif, Raja Jurdak +1

In this paper, we propose to utilise diffusion models for data augmentation in speech emotion recognition (SER). In particular, we present an effective approach to utilise improved…

cs.SD20236 cited

Emotions Beyond Words: Non-Speech Audio Emotion Recognition With Edge Computing

Ibrahim Malik, Siddique Latif, Sanaullah Manzoor +3

Non-speech emotion recognition has a wide range of applications including healthcare, crime control and rescue, and entertainment, to name a few. Providing these applications using…

cs.SD20231 cited

Lightweight Toxicity Detection in Spoken Language: A Transformer-based Approach for Edge Devices

Ahlam Husni Abu Nada, Siddique Latif, Junaid Qadir

Toxicity is a prevalent social behavior that involves the use of hate speech, offensive language, bullying, and abusive speech. While text-based approaches for toxicity detection a…

cs.SD20232 cited

Generative Emotional AI for Speech Emotion Recognition: The Case for Synthetic Emotional Speech Augmentation

Abdullah Shahid, Siddique Latif, Junaid Qadir

Despite advances in deep learning, current state-of-the-art speech emotion recognition (SER) systems still have poor performance due to a lack of speech emotion datasets. This pape…