3 papers
cs.AR2026
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
Chen Zhang, Qijun Zhang, Zhuoshan Zhou +10
Tensor parallelism (TP) in large-scale LLM inference and training introduces frequent collective operations that dominate inter-GPU communication. While in-switch computing, exempl…
cs.OS2025
CXLAimPod: CXL Memory is all you need in AI era
Yiwei Yang, Yusheng Zheng, Yiqi Chen +5
The proliferation of data-intensive applications, ranging from large language models to key-value stores, increasingly stresses memory systems with mixed read-write access patterns…
cs.AR2025
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
Yiqi Chen, Xiping Dong, Zhe Zhou +3
The Compute Express Link (CXL) technology facilitates the extension of CPU memory through byte-addressable SerDes links and cascaded switches, creating complex heterogeneous memory…