activity
20172022
most citedDynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

311 citations · 673 across the 44 of their papers we have counts for

collaborators

56 papers

cs.CL20221 cited

Rephrasing the Reference for Non-Autoregressive Machine Translation

Chenze Shao, Jinchao Zhang, Jie Zhou +1

Non-autoregressive neural machine translation (NAT) models suffer from the multi-modality problem that there may exist multiple possible translations of a source sentence, so the r…

cs.AI20223 cited

AutoCAD: Automatically Generating Counterfactuals for Mitigating Shortcut Learning

Jiaxin Wen, Yeshuang Zhu, Jinchao Zhang +2

Recent studies have shown the impressive efficacy of counterfactually augmented data (CAD) for reducing NLU models' reliance on spurious features and improving their generalizabili…

cs.CL20221 cited

MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction

Xiaozhi Wang, Yulin Chen, Ning Ding +9

The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two…

cs.CL2022

Counterfactual Data Augmentation via Perspective Transition for Open-Domain Dialogues

Jiao Ou, Jinchao Zhang, Yang Feng +1

The construction of open-domain dialogue systems requires high-quality dialogue datasets. The dialogue data admits a wide variety of responses for a given dialogue history, especia…

cs.CL2022

Exploring Mode Connectivity for Pre-trained Language Models

Yujia Qin, Cheng Qian, Jing Yi +6

Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP. From the perspective of parameter space, PLMs provide generic initialization, st…

cs.CL20222 cited

Different Tunes Played with Equal Skill: Exploring a Unified Optimization Subspace for Delta Tuning

Jing Yi, Weize Chen, Yujia Qin +6

Delta tuning (DET, also known as parameter-efficient tuning) is deemed as the new paradigm for using pre-trained language models (PLMs). Up to now, various DETs with distinct desig…