424 citations · 474 across the 24 of their papers we have counts for
3 papers · 1 filter
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Mostafa Elhoushi, Alex Pretko, Nolan Dey +6
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformer…
CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
Anisha Garg, Claire Zhang, Nishit Neema +3
Group-Relative Policy Optimization (GRPO) has emerged as the standard for training reasoning capabilities in large language models through reinforcement learning. By estimating adv…
BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model
Nolan Dey, Daria Soboleva, Faisal Al-Khateeb +11
We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from th…