28 citations · 41 across the 4 of their papers we have counts for
5 papers
Listen, Look and Deliberate: Visual context-aware speech recognition using pre-trained text-video representations
Shahram Ghorbani, Yashesh Gaur, Yu Shi +1
In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR). We explore…
Audio-visual Recognition of Overlapped speech for the LRS2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu +7
Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of…
Domain Expansion in DNN-based Acoustic Models for Robust Speech Recognition
Shahram Ghorbani, Soheil Khorram, John H. L. Hansen
Training acoustic models with sequentially incoming data -- while both leveraging new data and avoiding the forgetting effect-- is an essential obstacle to achieving human intellig…
Leveraging native language information for improved accented speech recognition
Shahram Ghorbani, John H. L. Hansen
Recognition of accented speech is a long-standing challenge for automatic speech recognition (ASR) systems, given the increasing worldwide population of bi-lingual speakers with En…
Advancing Multi-Accented LSTM-CTC Speech Recognition using a Domain Specific Student-Teacher Learning Paradigm
Shahram Ghorbani, Ahmet E. Bulut, John H. L. Hansen
Non-native speech causes automatic speech recognition systems to degrade in performance. Past strategies to address this challenge have considered model adaptation, accent classifi…