2 papers
cs.LG2026
Acceptance-Aware Draft Model Training for Speculative Decoding
Tianhua Xia, Mugilan Ganesan, Yifei Feng +3
Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to generate multiple candidate tokens that are verified by the target model…
cs.LG2025
MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
Mugilan Ganesan, Shane Segal, Ankur Aggarwal +3
Speculative decoding significantly accelerates language model inference by enabling a lightweight draft model to propose multiple tokens that a larger target model verifies simulta…