Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
Ryan Sun, Tianyi Zhou, Xun Chen +1
Large Language Models (LLMs) have become essential in advancing natural language processing (NLP) tasks, but their sequential token generation limits inference speed. Multi-Draft S…
cs.CL2024
BenTo: Benchmark Task Reduction with In-Context Transferability
Hongyu Zhao, Ming Li, Lichao Sun +1
Evaluating large language models (LLMs) is costly: it requires the generation and examination of LLM outputs on a large-scale benchmark of various tasks. This paper investigates ho…