1 paper · 1 filter
Yuheng Zhang, Yizhao Wang, Da Zhu +19
We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive p…