54 citations · 56 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 2 cited
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
Bowen Pan, Yikang Shen, Haokun Liu +5
Mixture-of-Experts (MoE) language models can reduce computational costs by 2-4 compared to dense models without sacrificing performance, making them more efficient in compu…
cs.SE2023★ 54 cited
SantaCoder: don't reach for the stars!
Loubna Ben Allal, Raymond Li, Denis Kocetkov +38
The BigCode project is an open-scientific collaboration working on the responsible development of large language models for code. This tech report describes the progress of the col…