2 papers
cs.CL2026
SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding
Yijun Lin, Jinhao Sheng, Qingyue Cai +1
Autoregressive language models suffer from high inference latency due to their sequential decoding nature. Speculative decoding (SD) mitigates this by employing a lightweight draft…
cs.DC2025
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
Jinhao Sheng, Zhiqing Tang, Jianxiong Guo +1
The growing demand for real-time processing tasks is driving the need for multi-model inference pipelines on edge devices. However, cost-effectively deploying these pipelines while…