activity
20192022
most citedCross-lingual Information Retrieval with BERT

22 citations · 27 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2022

Training Autoregressive Speech Recognition Models with Limited in-domain Supervision

Chak-Fai Li, Francis Keith, William Hartmann +1

Advances in self-supervised learning have significantly reduced the amount of transcribed audio required for training. However, the majority of work in this area is focused on read…

cs.CL20213 cited

Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated Pseudotranscripts

Chak-Fai Li, Francis Keith, William Hartmann +2

Sequence-to-sequence (seq2seq) models are competitive with hybrid models for automatic speech recognition (ASR) tasks when large amounts of training data are available. However, da…

cs.CL2021

Using heterogeneity in semi-supervised transcription hypotheses to improve code-switched speech recognition

Andrew Slottje, Shannon Wotherspoon, William Hartmann +2

Modeling code-switched speech is an important problem in automatic speech recognition (ASR). Labeled code-switched data are rare, so monolingual data are often used to model code-s…

cs.CV20202 cited

Learning from Noisy Labels with Noise Modeling Network

Zhuolin Jiang, Jan Silovsky, Man-Hung Siu +3

Multi-label image classification has generated significant interest in recent years and the performance of such systems often suffers from the not so infrequent occurrence of incor…

cs.IR202022 cited

Cross-lingual Information Retrieval with BERT

Zhuolin Jiang, Amro El-Jaroudi, William Hartmann +2

Multiple neural language models have been developed recently, e.g., BERT and XLNet, and achieved impressive results in various NLP tasks including sentence classification, question…

cs.LG2019

Towards a New Understanding of the Training of Neural Networks with Mislabeled Training Data

Herbert Gish, Jan Silovsky, Man-Ling Sung +3

We investigate the problem of machine learning with mislabeled training data. We try to make the effects of mislabeled training better understood through analysis of the basic mode…