137 citations · 137 across the 3 of their papers we have counts for
11 papers
Capitalization Normalization for Language Modeling with an Accurate and Efficient Hierarchical RNN Model
Hao Zhang, You-Chi Cheng, Shankar Kumar +3
Capitalization normalization (truecasing) is the task of restoring the correct case (uppercase or lowercase) of noisy text. We propose a fast, accurate and compact two-level hierar…
Scaling End-to-End Models for Large-Scale Multilingual ASR
Bo Li, Ruoming Pang, Tara N. Sainath +7
Building ASR models across many languages is a challenging multi-task learning problem due to large variations and heavily unbalanced data. Existing work has shown positive transfe…
Lookup-Table Recurrent Language Models for Long Tail Speech Recognition
W. Ronny Huang, Tara N. Sainath, Cal Peyser +3
We introduce Lookup-Table Language Models (LookupLM), a method for scaling up the size of RNN language models with only a constant increase in the floating point operations, by inc…
MetaPoison: Practical General-purpose Clean-label Data Poisoning
W. Ronny Huang, Jonas Geiping, Liam Fowl +2
Data poisoning -- the process by which an attacker takes control of a model by making imperceptible changes to a subset of the training data -- is an emerging threat in the context…
DeepErase: Weakly Supervised Ink Artifact Removal in Document Text Images
W. Ronny Huang, Yike Qi, Qianqian Li +1
Paper-intensive industries like insurance, law, and government have long leveraged optical character recognition (OCR) to automatically transcribe hordes of scanned documents into…
Deep k-NN Defense against Clean-label Data Poisoning Attacks
Neehar Peri, Neal Gupta, W. Ronny Huang +5
Targeted clean-label data poisoning is a type of adversarial attack on machine learning systems in which an adversary injects a few correctly-labeled, minimally-perturbed samples i…