papers
Publications (3)
cs.AR2025
Architectural and System Implications of CXL-enabled Tiered Memory
Yujie Yang, Lingfeng Xiang, Peiran Du +7
Memory disaggregation is an emerging technology that decouples memory from traditional memory buses, enabling independent scaling of compute and memory. Compute Express Link (CXL),…
cs.OS2024
Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
Lingfeng Xiang, Zhen Lin, Weishu Deng +4
With the advent of byte-addressable memory devices, such as CXL memory, persistent memory, and storage-class memory, tiered memory systems have become a reality. Page migration is…
cs.LG2025
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
Weishu Deng, Yujie Yang, Peiran Du +6
Scaling inference for large language models (LLMs) is increasingly constrained by limited GPU memory, especially due to growing key-value (KV) caches required for long-context gene…