Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
Jingwei Song, Wanyi Chen, Xinyuan Song +7
Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to propose tokens that are later verified by a stronger target model. While…
cs.DC2025
Lattica: A Decentralized Cross-NAT Communication Framework for Scalable AI Inference and Training
Ween Yang, Jason Liu, Suli Wang +4
The rapid expansion of distributed Artificial Intelligence (AI) workloads beyond centralized data centers creates a demand for new communication substrates. These substrates must o…