2 citations · 7 across the 6 of their papers we have counts for
6 papers
A Multi-Modal Multilingual Benchmark for Document Image Classification
Yoshinari Fujinuma, Siddharth Varia, Nishant Sankaran +3
Document image classification is different from plain-text document classification and consists of classifying a document by understanding the content and structure of documents su…
Characterizing and Measuring Linguistic Dataset Drift
Tyler A. Chang, Kishaloy Halder, Neha Anna John +4
NLP models often degrade in performance when real world data distributions differ markedly from training data. However, existing dataset drift metrics in NLP have generally not con…
Taxonomy Expansion for Named Entity Recognition
Karthikeyan K, Yogarshi Vyas, Jie Ma +7
Training a Named Entity Recognition (NER) model often involves fixing a taxonomy of entity types. However, requirements evolve and we might need the NER model to recognize addition…
Comparing Biases and the Impact of Multilingual Training across Multiple Languages
Sharon Levy, Neha Anna John, Ling Liu +6
Studies in bias and fairness in natural language processing have primarily examined social biases within a single language and/or across few attributes (e.g. gender, race). However…
Simple Yet Effective Synthetic Dataset Construction for Unsupervised Opinion Summarization
Ming Shen, Jie Ma, Shuai Wang +4
Opinion summarization provides an important solution for summarizing opinions expressed among a large number of reviews. However, generating aspect-specific and general summaries i…
Dynamic Benchmarking of Masked Language Models on Temporal Concept Drift with Multiple Views
Katerina Margatina, Shuai Wang, Yogarshi Vyas +3
Temporal concept drift refers to the problem of data changing over time. In NLP, that would entail that language (e.g. new expressions, meaning shifts) and factual knowledge (e.g.…