3 papers
cs.AR2026
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
Kangbo Bai, Zhantong Zhu, Yifan Ding +1
In large-scale distributed LLM training, communication between devices becomes the key performance bottleneck. Chiplet technology can integrate multiple dies into a package to scal…
cs.AR2025
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
Siyuan He, Zhantong Zhu, Yandong He +1
LLM inference on mobile devices faces extraneous challenges due to limited memory bandwidth and computational resources. To address these issues, speculative inference and processi…
cs.AR2025
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
Zhantong Zhu, Hongou Li, Wenjie Ren +4
With the rapid advent of generative models, efficiently deploying these models on specialized hardware has become critical. Tensor Processing Units (TPUs) are designed to accelerat…