2 papers
cs.CL2026
FlexDraft: Flexible Speculative Decoding via Attention Tuning and Bonus-Guided Calibration
Yaojie Zhang, Jianuo Huang, Junlong Ke +5
Speculative decoding accelerates memory-bound LLM inference without quality degradation by using a fast drafter to propose multiple candidate tokens and the target model to verify…
cs.CV2026
DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion
Yongji Long, Shijun Liang, Jintao Li +1
Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing critical long-range informa…