most citedLightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning

9 citations · 14 across the 3 of their papers we have counts for

collaborators

6 papers

cs.IT20205 cited

Stable Phase Retrieval from Locally Stable and Conditionally Connected Measurements

Cheng Cheng, Ingrid Daubechies, Nadav Dym +1

This paper is concerned with stable phase retrieval for a family of phase retrieval models we name "locally stable and conditionally connected" (LSCC) measurement schemes. For ever…

cs.CL20209 cited

LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning

Kaitao Song, Hao Sun, Xu Tan +4

While pre-training and fine-tuning, e.g., BERT~\citep{devlin2018bert}, GPT-2~\citep{radford2019language}, have achieved great success in language understanding and generation tasks…

cs.CL2020

MPNet: Masked and Permuted Pre-training for Language Understanding

Kaitao Song, Xu Tan, Tao Qin +2

BERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet intr…

math.AP2020

Non-Convex Planar Harmonic Maps

Shahar Z. Kovalsky, Noam Aigerman, Ingrid Daubechies +3

We formulate a novel characterization of a family of invertible maps between two-dimensional domains. Our work follows two classic results: The Radó-Kneser-Choquet (RKC) theorem, w…

cs.CL2018

Hybrid Self-Attention Network for Machine Translation

Kaitao Song, Xu Tan, Furong Peng +1

The encoder-decoder is the typical framework for Neural Machine Translation (NMT), and different structures have been developed for improving the translation performance. Transform…

cs.CV2018

Stop memorizing: A data-dependent regularization framework for intrinsic pattern learning

Wei Zhu, Qiang Qiu, Bao Wang +3

Deep neural networks (DNNs) typically have enough capacity to fit random data by brute force even when conventional data-dependent regularizations focusing on the geometry of the f…