From the 1 of 4 linked papers with an AI index.
4 papers
C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems
Jiayi Li, Di Wu, Qingxu Li +10
The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communicati…
CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation
Yuanpeng Zhang, YuXuan Wu, Yitong Xiao +6
The paper introduces CODA, a hardware-software co-designed architecture that separates compute and cache operations for edge video diffusion models, using near‑memory processing to…
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
Zhao Wang, Yiqi Chen, Cong Li +5
Transaction processing systems are the crux for modern data-center applications, yet current multi-node systems are slow due to network overheads. This paper advocates for Compute…
Unicorn: Unified Neural Image Compression with One Number Reconstruction
Qi Zheng, Haozhi Wang, Zihao Liu +8
Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit im…