activity
20162023
most citedA cross-corpus study on speech emotion recognition

31 citations · 66 across the 20 of their papers we have counts for

collaborators
Showing cs.CLShow all

17 papers · 1 filter

cs.CL2024

Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis

Cong-Thanh Do, Shuhei Imai, Rama Doddipatla +1

This paper investigates the use of unsupervised text-to-speech synthesis (TTS) as a data augmentation method to improve accented speech recognition. TTS systems are trained with a…

cs.CL2024

LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks

Amit Meghanani, Thomas Hain

Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representati…

cs.CL2024

Automatic Speech Recognition System-Independent Word Error Rate Estimation

Chanho Park, Mingjie Chen, Thomas Hain

Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition (ASR) systems. In many applications, it is of interest to…

cs.CL2024

Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations

Amit Meghanani, Thomas Hain

Acoustic word embeddings (AWEs) are vector representations of spoken words. An effective method for obtaining AWEs is the Correspondence Auto-Encoder (CAE). In the past, the CAE me…

cs.CL2024

SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations

Amit Meghanani, Thomas Hain

There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. Thes…

cs.CL2023

MUST: A Multilingual Student-Teacher Learning approach for low-resource speech recognition

Muhammad Umar Farooq, Rehan Ahmad, Thomas Hain

Student-teacher learning or knowledge distillation (KD) has been previously used to address data scarcity issue for training of speech recognition (ASR) systems. However, a limitat…