2 papers
cs.LG2025
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
Weishu Deng, Yujie Yang, Peiran Du +6
Scaling inference for large language models (LLMs) is increasingly constrained by limited GPU memory, especially due to growing key-value (KV) caches required for long-context gene…
cs.AR2025
Architectural and System Implications of CXL-enabled Tiered Memory
Yujie Yang, Lingfeng Xiang, Peiran Du +7
Memory disaggregation is an emerging technology that decouples memory from traditional memory buses, enabling independent scaling of compute and memory. Compute Express Link (CXL),…