2 papers
cs.LG2025
Draft, Verify, and Improve: Toward Training-Aware Speculative Decoding
Shrenik Bhansali, Larry Heck
Autoregressive (AR) decoding is a major latency bottleneck for large language models. Speculative decoding (SD) accelerates AR by letting a drafter propose multi-token blocks that…
cs.CL2024
LEGO: Language Model Building Blocks
Shrenik Bhansali, Alwin Jin, Tyler Lizzo +1
Large language models (LLMs) are essential in natural language processing (NLP) but are costly in data collection, pre-training, fine-tuning, and inference. Task-specific small lan…