Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Cross-Attention Speculative Decoding
Wei Zhong, Manasa Bharadwaj, Yixiao Wang +2
Speculative decoding (SD) is a widely adopted approach for accelerating inference in large language models (LLMs), particularly when the draft and target models are well aligned. H…
cs.CL2024
S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs
Wei Zhong, Manasa Bharadwaj
Speculative decoding (SD) has attracted a significant amount of research attention due to the substantial speedup it can achieve for LLM inference. However, despite the high speedu…