2 papers
cs.NI2026
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
Zengzipeng Tang, Yuxuan Sun, Wei Chen +2
The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference. As a promising architecture, Specula…
cs.NI2026
Hierarchical Online-Scheduling for Energy-Efficient Split Inference with Progressive Transmission
Zengzipeng Tang, Yuxuan Sun, Wei Chen +3
Device-edge collaborative inference with Deep Neural Networks (DNNs) faces fundamental trade-offs among accuracy, latency and energy consumption. Current scheduling exhibits two dr…