4 papers
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
Cong Li, Yihan Yin, Chenhao Xue +7
Large language models (LLMs) have been widely deployed for online generative services, where numerous LLM instances jointly handle workloads with fluctuating request arrival rates…
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
Yiqi Chen, Xiping Dong, Zhe Zhou +3
The Compute Express Link (CXL) technology facilitates the extension of CPU memory through byte-addressable SerDes links and cascaded switches, creating complex heterogeneous memory…
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
Zhao Wang, Yiqi Chen, Cong Li +5
Transaction processing systems are the crux for modern data-center applications, yet current multi-node systems are slow due to network overheads. This paper advocates for Compute…
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
Jingchen Zhu, Chenhao Xue, Yiqi Chen +13
The emergence of the large language model~(LLM) poses an exponential growth of demand for computation throughput, memory capacity, and communication bandwidth. Such a demand growth…