papers

Publications (21)

eess.AS2023

SLICER: Learning universal audio representations using low-resource self-supervised pre-training

Ashish Seth, Sreyan Ghosh, S. Umesh +1

We present a new Self-Supervised Learning (SSL) approach to pre-train encoders on unlabeled audio data that reduces the need for large amounts of labeled data for audio and speech…

cs.CL2022

DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances

Sreyan Ghosh, Samden Lepcha, S Sakshi +2

Toxic speech, also known as hate speech, is regarded as one of the crucial issues plaguing online social media today. Most recent work on toxic speech detection is constrained to t…

cs.SD2023

DECAR: Deep Clustering for learning general-purpose Audio Representations

Sreyan Ghosh, Sandesh V Katta, Ashish Seth +1

We introduce DECAR, a self-supervised pre-training approach for learning general-purpose audio representations. Our system is based on clustering: it utilizes an offline clustering…

eess.AS2023

FusDom: Combining In-Domain and Out-of-Domain Knowledge for Continuous Self-Supervised Learning

Ashish Seth, Sreyan Ghosh, S. Umesh +1

Continued pre-training (CP) offers multiple advantages, like target domain adaptation and the potential to exploit the continuous stream of unlabeled data available online. However…

eess.AS2023

Channel-Aware Pretraining of Joint Encoder-Decoder Self-Supervised Model for Telephonic-Speech ASR

Vrunda N. Sukhadia, A. Arunkumar, S. Umesh

This paper proposes a novel technique to obtain better downstream ASR performance from a joint encoder-decoder self-supervised model when trained with speech pooled from two differ…

eess.AS2023

data2vec-aqc: Search for the right Teaching Assistant in the Teacher-Student training setup

Vasista Sai Lodagala, Sreyan Ghosh, S. Umesh

In this paper, we propose a new Self-Supervised Learning (SSL) algorithm called data2vec-aqc, for speech representation learning from unlabeled speech data. Our goal is to improve…

cs.CL2023

The Tag-Team Approach: Leveraging CLS and Language Tagging for Enhancing Multilingual ASR

Kaousheik Jayakumar, Vrunda N. Sukhadia, A Arunkumar +1

Building a multilingual Automated Speech Recognition (ASR) system in a linguistically diverse country like India can be a challenging task due to the differences in scripts and the…

eess.AS2023

Domain Adaptation of low-resource Target-Domain models using well-trained ASR Conformer Models

Vrunda N. Sukhadia, S. Umesh

In this paper, we investigate domain adaptation for low-resource Automatic Speech Recognition (ASR) of target-domain data, when a well-trained ASR model trained with a large datase…

cs.CL2023

Analyzing the factors affecting usefulness of Self-Supervised Pre-trained Representations for Speech Recognition

Ashish Seth, Lodagala V S V Durga Prasad, Sreyan Ghosh +1

Self-supervised learning (SSL) to learn high-level speech representations has been a popular approach to building Automatic Speech Recognition (ASR) systems in low-resource setting…

cs.CL2022

Span Classification with Structured Information for Disfluency Detection in Spoken Utterances

Sreyan Ghosh, Sonal Kumar, Yaman Kumar Singla +2

Existing approaches in disfluency detection focus on solving a token-level classification task for identifying and removing disfluencies in text. Moreover, most works focus on leve…

cs.SD2022

DeLoRes: Decorrelating Latent Spaces for Low-Resource Audio Representation Learning

Sreyan Ghosh, Ashish Seth, and Deepak Mittal +2

Inspired by the recent progress in self-supervised learning for computer vision, in this paper we introduce DeLoRes, a new general-purpose audio representation learning approach. O…

eess.AS2023

UNFUSED: UNsupervised Finetuning Using SElf supervised Distillation

Ashish Seth, Sreyan Ghosh, S. Umesh +1

In this paper, we introduce UnFuSeD, a novel approach to leverage self-supervised learning and reduce the need for large amounts of labeled data for audio classification. Unlike pr…

eess.AS2021

Investigation of Speaker-adaptation methods in Transformer based ASR

Vishwas M. Shetty, Metilda Sagaya Mary N J, S. Umesh

End-to-end models are fast replacing the conventional hybrid models in automatic speech recognition. Transformer, a sequence-to-sequence model, based on self-attention popularly us…

cs.CL2025

Building Robust and Scalable Multilingual ASR for Indian Languages

Arjun Gangwar, Kaousheik Jayakumar, S. Umesh

This paper describes the systems developed by SPRING Lab, Indian Institute of Technology Madras, for the ASRU MADASR 2.0 challenge. The systems developed focuses on adapting ASR sy…

cs.CL2023

PADA: Pruning Assisted Domain Adaptation for Self-Supervised Speech Representations

Lodagala V S V Durga Prasad, Sreyan Ghosh, S. Umesh

While self-supervised speech representation learning (SSL) models serve a variety of downstream tasks, these models have been observed to overfit to the domain from which the unlab…

cs.CL2022

Investigation of Ensemble features of Self-Supervised Pretrained Models for Automatic Speech Recognition

A Arunkumar, Vrunda N Sukhadia, S. Umesh

Self-supervised learning (SSL) based models have been shown to generate powerful representations that can be used to improve the performance of downstream speech tasks. Several sta…

cs.LG2013

Modified SPLICE and its Extension to Non-Stereo Data for Noise Robust Speech Recognition

D. S. Pavan Kumar, N. Vishnu Prasad, Vikas Joshi +1

In this paper, a modification to the training process of the popular SPLICE algorithm has been proposed for noise robust speech recognition. The modification is based on feature co…

eess.AS2023

MAST: Multiscale Audio Spectrogram Transformers

Sreyan Ghosh, Ashish Seth, S. Umesh +1

We present Multiscale Audio Spectrogram Transformer (MAST) for audio classification, which brings the concept of multiscale feature hierarchies to the Audio Spectrogram Transformer…

cs.SD2025

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

Advait Joglekar, Divyanshu Singh, Rooshil Rohit Bhatia +1

Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectur…

cs.CL2023

CCC-wav2vec 2.0: Clustering aided Cross Contrastive Self-supervised learning of speech representations

Vasista Sai Lodagala, Sreyan Ghosh, S. Umesh

While Self-Supervised Learning has helped reap the benefit of the scale from the available unlabeled data, the learning paradigms are continuously being bettered. We present a new…

eess.AS2023

Stable Distillation: Regularizing Continued Pre-training for Low-Resource Automatic Speech Recognition

Ashish Seth, Sreyan Ghosh, S. Umesh +1

Continued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain has shown to be extremely effective for low-resource Automatic Speech Recognition…