activity
20162024
most citedLIMA: Less Is More for Alignment

128 citations · 168 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL202322 cited

The Temporal Structure of Language Processing in the Human Brain Corresponds to The Layered Hierarchy of Deep Language Models

Ariel Goldstein, Eric Ham, Mariano Schain +17

Deep Language Models (DLMs) provide a novel computational paradigm for understanding the mechanisms of natural language processing in the human brain. Unlike traditional psycholing…

cs.CL2023128 cited

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu, Puxin Xu +12

Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and re…

cs.CL20237 cited

Scaling Laws for Generative Mixed-Modal Language Models

Armen Aghajanyan, Lili Yu, Alexis Conneau +7

Generative language models define distributions over sequences of tokens that can represent essentially any combination of data modalities (e.g., any permutation of image tokens fr…

cs.CL2021

Simple Local Attentions Remain Competitive for Long-Context Tasks

Wenhan Xiong, Barlas Oğuz, Anchit Gupta +5

Many NLP tasks require processing long contexts beyond the length limit of pretrained models. In order to scale these models to longer text sequences, many efficient long-range att…

cs.CL20162 cited

A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence Alignments

Omer Levy, Anders Søgaard, Yoav Goldberg

While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague. We observe that whethe…