11 citations · 13 across the 13 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge
Haotian Zheng, Zhanwei Wang, Mingyao Cui +3
Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment t…
cs.DC2026
SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks
Zhanwei Wang, Huiling Yang, Min Sheng +2
Leveraging continuous solar energy harvesting at high efficiency, space data centers are envisioned as a promising platform for executing energy-intensive large language models (LL…