24 citations · 28 across the 9 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Mostafa Elhoushi, Alex Pretko, Nolan Dey +6
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformer…
cs.AI2023★ 2 cited
BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model
Nolan Dey, Daria Soboleva, Faisal Al-Khateeb +11
We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from th…