2 papers
cs.LG2026
Attention Drift: What Autoregressive Speculative Decoding Models Learn
DoÄaç Eldenk, Payal Mohapatra, Yigitcan Comlek +3
Speculative decoding accelerates LLM inference by drafting future tokens with a small model, but drafter models degrade sharply under template perturbation and long-context inputs.…
cs.LG2026
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
Yueyuan Sui, Payal Mohapatra, DoÄaç Eldenk +5
Edge devices increasingly run multimodal sensing pipelines that must remain accurate despite fluctuating power budgets and unpredictable sensor dropout. Existing pruning methods fa…