4 papers
Efficient infusion of self-supervised representations in Automatic Speech Recognition
Darshan Prabhu, Sai Ganesh Mirishkar, Pankaj Wasnik
Self-supervised learned (SSL) models such as Wav2vec and HuBERT yield state-of-the-art results on speech-related tasks. Given the effectiveness of such models, it is advantageous t…
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning
Shivam Ratnakant Mhaskar, Nirmesh J. Shah, Mohammadi Zaki +3
Traditional Automatic Video Dubbing (AVD) pipeline consists of three key modules, namely, Automatic Speech Recognition (ASR), Neural Machine Translation (NMT), and Text-to-Speech (…
Fiducial Focus Augmentation for Facial Landmark Detection
Purbayan Kar, Vishal Chudasama, Naoyuki Onoe +2
Deep learning methods have led to significant improvements in the performance on the facial landmark detection (FLD) task. However, detecting landmarks in challenging settings, suc…
Revisiting Class Imbalance for End-to-end Semi-Supervised Object Detection
Purbayan Kar, Vishal Chudasama, Naoyuki Onoe +1
Semi-supervised object detection (SSOD) has made significant progress with the development of pseudo-label-based end-to-end methods. However, many of these methods face challenges…