activity
20182021
most citedHierarchical Topic Mining via Joint Spherical Tree and Text Embedding

59 citations · 158 across the 8 of their papers we have counts for

collaborators

16 papers

cs.CL20222 cited

Pretraining Text Encoders with Adversarial Mixture of Training Signal Generators

Yu Meng, Chenyan Xiong, Payal Bajaj +4

We present a new framework AMOS that pretrains text encoders with an Adversarial learning curriculum via a Mixture Of Signals from multiple auxiliary generators. Following ELECTRA-…

cs.CL202250 cited

Topic Discovery via Latent Space Clustering of Pretrained Language Model Representations

Yu Meng, Yunyi Zhang, Jiaxin Huang +2

Topic models have been the prominent tools for automatic topic discovery from text corpora. Despite their effectiveness, topic models suffer from several limitations including the…

cs.CL20211 cited

Fine-Grained Opinion Summarization with Minimal Supervision

Suyu Ge, Jiaxin Huang, Yu Meng +2

Opinion summarization aims to profile a target by extracting opinions from multiple documents. Most existing work approaches the task in a semi-supervised manner due to the difficu…

cs.CL20211 cited

Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-Training

Yu Meng, Yunyi Zhang, Jiaxin Huang +4

We study the problem of training named entity recognition (NER) models using only distantly-labeled data, which can be automatically obtained by matching entity mentions in the raw…

cs.CL202118 cited

UCPhrase: Unsupervised Context-aware Quality Phrase Tagging

Xiaotao Gu, Zihan Wang, Zhenyu Bi +4

Identifying and understanding quality phrases from context is a fundamental task in text mining. The most challenging part of this task arguably lies in uncommon, emerging, and dom…

cs.CL2021

COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining

Yu Meng, Chenyan Xiong, Payal Bajaj +4

We present a self-supervised learning framework, COCO-LM, that pretrains Language Models by COrrecting and COntrasting corrupted text sequences. Following ELECTRA-style pretraining…