2 papers
cs.AR2026
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
Kangbo Bai, Zhantong Zhu, Yifan Ding +1
In large-scale distributed LLM training, communication between devices becomes the key performance bottleneck. Chiplet technology can integrate multiple dies into a package to scal…
cs.AR2026
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
Chenhao Xue, Yukun Wang, An Guo +11
SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip d…