6 citations · 10 across the 33 of their papers we have counts for
3 papers · 1 filter
SpliceOut: A Simple and Efficient Audio Augmentation Method
Arjit Jain, Pranay Reddy Samala, Deepak Mittal +2
Time masking has become a de facto augmentation technique for speech and audio tasks, including automatic speech recognition (ASR) and audio classification, most notably as a part…
Cross-Modal learning for Audio-Visual Video Parsing
Jatin Lamba, Abhishek, Jayaprakash Akula +3
In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The propose…
Error-driven Fixed-Budget ASR Personalization for Accented Speakers
Abhijeet Awasthi, Aman Kansal, Sunita Sarawagi +1
We consider the task of personalizing ASR models while being constrained by a fixed budget on recording speaker-specific utterances. Given a speaker and an ASR model, we propose a…