2 papers
cs.DC2025
nncase: An End-to-End Compiler for Efficient LLM Deployment on Heterogeneous Storage Architectures
Hui Guo, Qihang Zheng, Chenghai Huo +3
The efficient deployment of large language models (LLMs) is hindered by memory architecture heterogeneity, where traditional compilers suffer from fragmented workflows and high ada…
cs.NI2025
Arcturus: A Cloud Overlay Network for Global Accelerator with Enhanced Performance and Stability
Matthew Yang Liu, Chuang Chen, Pengcheng Lv +6
Global Accelerator (GA) services play a vital role in ensuring low-latency, high-reliability communication for real-time interactive applications. However, existing GA offerings ar…