3 papers
cs.DC2026
NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching
Nobel Dhar, Md Romyull Islam, Xuechen Zhang +4
Deploying large language models on edge devices is increasingly limited by a widening gap between model size and available memory. Existing approaches such as quantization, smaller…
cs.AI2025
Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
Zijun Wang, Yijiahao Qi, Hanqiu Chen +5
Mixture-of-Agents (MoA) inference can suffer from dense inter-agent communication and low hardware utilization, which jointly inflate serving latency. We present a serving design t…
cs.AR2024
ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture Model
Hanqiu Chen, Yitu Wang, Luis Vitorio Cargnini +7
Compute Express Link (CXL) emerges as a solution for wide gap between computational speed and data communication rates among host and multiple devices. It fosters a unified and coh…