24 citations · 27 across the 3 of their papers we have counts for
3 papers
cs.CL2023★ 1 cited
Position Interpolation Improves ALiBi Extrapolation
Faisal Al-Khateeb, Nolan Dey, Daria Soboleva +1
Linear position interpolation helps pre-trained models using rotary position embeddings (RoPE) to extrapolate to longer sequence lengths. We propose using linear position interpola…
cs.AI2023★ 2 cited
BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model
Nolan Dey, Daria Soboleva, Faisal Al-Khateeb +11
We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from th…
cs.LG2023★ 24 cited
Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
Nolan Dey, Gurpreet Gosal, Zhiming +6
We study recent research advances that improve large language models through efficient pre-training and scaling, and open datasets and tools. We combine these advances to introduce…