5 papers · 1 filter
A Fusion-Aware Two-Stage Framework for Mispronunciation Detection and Diagnosis in Low-Resource Modern Standard Arabic
Jing Yang, Shuqing Zhang, Yongyi Deng +5
Accurate phoneme recognition is pivotal for mispronunciation detection and diagnosis (MDD) in modern standard Arabic (MSA), yet remains constrained by data scarcity and the synthet…
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
Xin Guo, Chunrui Zhao, Hong Jia +4
Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant…
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
Trevor Adelson, Vidhyasaharan Sethu, Ting Dang
Deep learning dominates speech processing but relies on massive datasets, global backpropagation-guided weight updates, and produces entangled representations. Assembly Calculus (A…
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
Wenda Zhang, Hongyu Jin, Siyi Wang +2
Speech Emotion Recognition models typically use single categorical labels, overlooking the inherent ambiguity of human emotions. Ambiguous Emotion Recognition addresses this by rep…
Speech Emotion Recognition Via CNN-Transformer and Multidimensional Attention Mechanism
Xiaoyu Tang, Yixin Lin, Ting Dang +2
Speech Emotion Recognition (SER) is crucial in human-machine interactions. Mainstream approaches utilize Convolutional Neural Networks or Recurrent Neural Networks to learn local e…