activity
20182026
most citedLarge Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias

72 citations · 362 across the 34 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

cs.CL2021

MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information

Yu Zhang, Shweta Garg, Yu Meng +2

We study the problem of weakly supervised text classification, which aims to classify text documents into a set of pre-defined categories with category surface names only and witho…

cs.CL2021★ 1 cited

Fine-Grained Opinion Summarization with Minimal Supervision

Suyu Ge, Jiaxin Huang, Yu Meng +2

Opinion summarization aims to profile a target by extracting opinions from multiple documents. Most existing work approaches the task in a semi-supervised manner due to the difficu…

cs.CL2021★ 1 cited

Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-Training

Yu Meng, Yunyi Zhang, Jiaxin Huang +4

We study the problem of training named entity recognition (NER) models using only distantly-labeled data, which can be automatically obtained by matching entity mentions in the raw…

cs.CL2021★ 18 cited

UCPhrase: Unsupervised Context-aware Quality Phrase Tagging

Xiaotao Gu, Zihan Wang, Zhenyu Bi +4

Identifying and understanding quality phrases from context is a fundamental task in text mining. The most challenging part of this task arguably lies in uncommon, emerging, and dom…

cs.CL2021

COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining

Yu Meng, Chenyan Xiong, Payal Bajaj +4

We present a self-supervised learning framework, COCO-LM, that pretrains Language Models by COrrecting and COntrasting corrupted text sequences. Following ELECTRA-style pretraining…