Showing cs.CLShow all
2 papers · 1 filter
cs.CL2020
hinglishNorm -- A Corpus of Hindi-English Code Mixed Sentences for Text Normalization
Piyush Makhija, Ankit Kumar, Anuj Gupta
We present hinglishNorm -- a human annotated corpus of Hindi-English code-mixed sentences for text normalization task. Each sentence in the corpus is aligned to its corresponding h…
cs.CL2020
Noisy Text Data: Achilles' Heel of BERT
Ankit Kumar, Piyush Makhija, Anuj Gupta
Owing to the phenomenal success of BERT on various NLP tasks and benchmark datasets, industry practitioners are actively experimenting with fine-tuning BERT to build NLP applicatio…