24 citations · 26 across the 2 of their papers we have counts for
2 papers
cs.AI2023★ 2 cited
BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model
Nolan Dey, Daria Soboleva, Faisal Al-Khateeb +11
We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from th…
cs.LG2023★ 24 cited
Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
Nolan Dey, Gurpreet Gosal, Zhiming +6
We study recent research advances that improve large language models through efficient pre-training and scaling, and open datasets and tools. We combine these advances to introduce…