2 papers
cs.AR2026
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
Rui Xie, Asad Ul Haq, Yunhua Fang +4
LLM inference is increasingly limited by memory bandwidth, and the bottleneck worsens at long context as the KV cache grows. CXL memory adds capacity to offload weights and KV, but…
cs.AR2025
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
Rui Xie, Asad Ul Haq, Linsen Ma +4
The efficiency of Large Language Model~(LLM) inference is often constrained by substantial memory bandwidth and capacity demands. Existing techniques, such as pruning, quantization…