2 papers
cs.DC2026
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
Yunhe Han, Yunqi Gao, Bing Hu +4
Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness,…
cs.CL2025
aiXcoder-7B: A Lightweight and Effective Large Language Model for Code Processing
Siyuan Jiang, Jia Li, He Zong +11
Large Language Models (LLMs) have been widely used in code completion, and researchers are focusing on scaling up LLMs to improve their accuracy. However, larger LLMs have lower in…