activity
20102026
most citedTransition-Based Dependency Parsing with Stack Long Short-Term Memory

526 citations · 1.3k across the 34 of their papers we have counts for

collaborators
Showing 2022Show all

5 papers · 1 filter

cs.CL2022

Domain Mismatch Doesn't Always Prevent Cross-Lingual Transfer Learning

Daniel Edmiston, Phillip Keung, Noah A. Smith

Cross-lingual transfer learning without labeled target language data or parallel text has been surprisingly effective in zero-shot cross-lingual classification, question answering,…

cs.CL20222 cited

How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers

Michael Hassid, Hao Peng, Daniel Rotem +4

The attention mechanism is considered the backbone of the widely-used Transformer architecture. It contextualizes the input by computing input-specific attention matrices. We find…

cs.CL20221 cited

Modeling Context With Linear Attention for Scalable Document-Level Translation

Zhaofeng Wu, Hao Peng, Nikolaos Pappas +1

Document-level machine translation leverages inter-sentence dependencies to produce more coherent and consistent translations. However, these models, predominantly based on transfo…

cs.CL202264 cited

Selective Annotation Makes Language Models Better Few-Shot Learners

Hongjin Su, Jungo Kasai, Chen Henry Wu +8

Many recent approaches to natural language tasks are built on the remarkable abilities of large language models. Large language models can perform in-context learning, where they l…

cs.CL20226 cited

Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Suchin Gururangan, Dallas Card, Sarah K. Dreier +5

Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, an…