2 papers
cs.CL2024
FastDraft: How to Train Your Draft
Ofir Zafrir, Igor Margulis, Dorin Shteyman +2
Speculative Decoding has gained popularity as an effective technique for accelerating the auto-regressive inference process of Large Language Models. However, Speculative Decoding…
cs.CL2022
Fast DistilBERT on CPUs
Haihao Shen, Ofir Zafrir, Bo Dong +7
Transformer-based language models have become the standard approach to solving natural language processing tasks. However, industry adoption usually requires the maximum throughput…