2 papers
cs.AR2026
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
Chenyang Ai, Yixing Zhang, Haoran Wu +3
Large language model (LLM) decoding is a major inference bottleneck because its low arithmetic intensity makes performance highly sensitive to memory bandwidth. 3D-stacked near-mem…
cs.AR2024
GTA: a new General Tensor Accelerator with Better Area Efficiency and Data Reuse
Chenyang Ai, Lechuan Zhao, Zhijie Huang +3
Recently, tensor algebra have witnessed significant applications across various domains. Each operator in tensor algebra features different computational workload and precision. Ho…