2 papers
cs.OS2026
GNStor: Design of GPU-Native High-Performance Remote All-Flash Array
Shushu Yi, Wenbo Wu, Guoci Chen +6
GPU has become the leading computing device for a wide range of data-intensive applications, which tightly collaborates with remote all-flash array (AFA) to accommodate ever-expand…
cs.DC2026
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
Lingfeng Tang, Daoping Zhang, Junjie Chen +6
Host-GPU data movement has become a latency-critical bottleneck in LLM serving, surfacing in common paths such as model-weight movement and KV cache offload/fetch. Today, each host…