49 citations · 49 across the 1 of their papers we have counts for
4 papers
Luna: Linear Unified Nested Attention
Xuezhe Ma, Xiang Kong, Sinong Wang +4
The quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Lun…
Entailment as Few-Shot Learner
Sinong Wang, Han Fang, Madian Khabsa +2
Large pre-trained language models (LMs) have demonstrated remarkable ability as few-shot learners. However, their success hinges largely on scaling model parameters to a degree tha…
On Unifying Misinformation Detection
Nayeon Lee, Belinda Z. Li, Sinong Wang +4
In this paper, we introduce UnifiedM2, a general-purpose misinformation model that jointly models multiple domains of misinformation with a single, unified setup. The model is trai…
On the Influence of Masking Policies in Intermediate Pre-training
Qinyuan Ye, Belinda Z. Li, Sinong Wang +5
Current NLP models are predominantly trained through a two-stage "pre-train then fine-tune" pipeline. Prior work has shown that inserting an intermediate pre-training stage, using…