activity
20192022
most citedCPM: A Large-scale Generative Chinese Pre-trained Language Model

22 citations · 48 across the 6 of their papers we have counts for

collaborators

11 papers

cs.CL20222 cited

Finding Skill Neurons in Pre-trained Transformer-based Language Models

Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang +3

Transformer-based pre-trained language models have demonstrated superior performance on various natural language processing tasks. However, it remains unclear how the skills requir…

cs.CL20221 cited

MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction

Xiaozhi Wang, Yulin Chen, Ning Ding +9

The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two…

cs.CL20222 cited

COPEN: Probing Conceptual Knowledge in Pre-trained Language Models

Hao Peng, Xiaozhi Wang, Shengding Hu +5

Conceptual knowledge is fundamental to human cognition and knowledge bases. However, existing knowledge probing works only focus on evaluating factual knowledge of pre-trained lang…

cs.CL2022

LEVEN: A Large-Scale Chinese Legal Event Detection Dataset

Feng Yao, Chaojun Xiao, Xiaozhi Wang +7

Recognizing facts is the most fundamental step in making judgments, hence detecting events in the legal documents is important to legal case analysis tasks. However, existing Legal…

cs.CL202221 cited

Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models

Ning Ding, Yujia Qin, Guang Yang +17

Despite the success, the process of fine-tuning large-scale PLMs brings prohibitive adaptation costs. In fact, fine-tuning all the parameters of a colossal model and retaining sepa…

cs.CL2021

CLEVE: Contrastive Pre-training for Event Extraction

Ziqi Wang, Xiaozhi Wang, Xu Han +6

Event extraction (EE) has considerably benefited from pre-trained language models (PLMs) by fine-tuning. However, existing pre-training methods have not involved modeling event cha…