1 paper
Aleksander Lorenc, Frédéric Berdoz, Joël Mathys +1
Improving the inference efficiency of autoregressive transformers typically means reducing FLOPs per token, usually through approximations that degrade model quality. We introduce…