1 paper
Kei-Sing Ng, Qingchen Wang
The success of large-scale language models like GPT can be attributed to their ability to efficiently predict the next token in a sequence. However, these models rely on constant c…