activity
20172022
most citedEmpower Sequence Labeling with Task-Aware Neural Language Model

151 citations · 258 across the 21 of their papers we have counts for

collaborators

34 papers

cs.CV2022

MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding

Zilong Wang, Jiuxiang Gu, Chris Tensmeyer +5

Document images are a ubiquitous source of data where the text is organized in a complex hierarchical structure ranging from fine granularity (e.g., words), medium granularity (e.g…

cs.CL2022

Progressive Sentiment Analysis for Code-Switched Text Data

Sudhanshu Ranjan, Dheeraj Mekala, Jingbo Shang

Multilingual transformer language models have recently attracted much attention from researchers and are used in cross-lingual transfer learning for many NLP tasks such as text cla…

cs.CL202221 cited

OA-Mine: Open-World Attribute Mining for E-Commerce Products with Weak Supervision

Xinyang Zhang, Chenwei Zhang, Xian Li +4

Automatic extraction of product attributes from their textual descriptions is essential for online shopper experience. One inherent challenge of this task is the emerging nature of…

cs.CL2022

Towards Few-shot Entity Recognition in Document Images: A Label-aware Sequence-to-Sequence Framework

Zilong Wang, Jingbo Shang

Entity recognition is a fundamental task in understanding document images. Traditional sequence labeling frameworks treat the entity types as class IDs and rely on extensive data a…

cs.CL2022

UCTopic: Unsupervised Contrastive Learning for Phrase Representations and Topic Mining

Jiacheng Li, Jingbo Shang, Julian McAuley

High-quality phrase representations are essential to finding topics and related terms in documents (a.k.a. topic mining). Existing phrase representation learning methods either sim…

cs.CL20211 cited

Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated Data

Dheeraj Mekala, Varun Gangal, Jingbo Shang

Existing text classification methods mainly focus on a fixed label set, whereas many real-world applications require extending to new fine-grained classes as the number of samples…