31 citations · 130 across the 14 of their papers we have counts for
4 papers · 1 filter
Improving a Named Entity Recognizer Trained on Noisy Data with a Few Clean Instances
Zhendong Chu, Ruiyi Zhang, Tong Yu +4
To achieve state-of-the-art performance, one still needs to train NER models on large-scale, high-quality annotated data, an asset that is both costly and time-intensive to accumul…
Learning the Visualness of Text Using Large Vision-Language Models
Gaurav Verma, Ryan A. Rossi, Christopher Tensmeyer +2
Visual text evokes an image in a person's mind, while non-visual text fails to do so. A method to automatically detect visualness in text will enable text-to-image retrieval and ge…
Unified Pretraining Framework for Document Understanding
Jiuxiang Gu, Jason Kuen, Vlad I. Morariu +5
Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabel…
Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU Models
Mengnan Du, Varun Manjunatha, Rajiv Jain +5
Recent studies indicate that NLU models are prone to rely on shortcut features for prediction, without achieving true language understanding. As a result, these models fail to gene…