Publications (6)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
Pingcheng Dong, Yonghao Tan, Xuejiao Liu +13
This work presents a 55nm speculative decoding-based LLM accelerator with bumping-based face-to-face ReRAM-on-logic stacking technology. It features a local rotation unit for outli…
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
Huanyu Qu, Weihao Zhang, Junfeng Lin +4
To efficiently support large-scale NNs, multi-level hardware, leveraging advanced integration and interconnection technologies, has emerged as a promising solution to counter the s…
CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
Hongyi Li, Songchen Ma, Huanyu Qu +5
The rapid advancement of Large Language Models (LLMs) has revolutionized various aspects of human life, yet their immense computational and energy demands pose significant challeng…
General-purpose Dataflow Model with Neuromorphic Primitives
Weihao Zhang, Yu Du, Hongyi Li +2
Neuromorphic computing exhibits great potential to provide high-performance benefits in various applications beyond neural networks. However, a general-purpose program execution mo…
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
Songchen Ma, Hongyi Li, Weihao Zhang +8
Mixture-of-Experts is a promising approach for edge AI with low-batch inference. Yet, on-device deployments often face limited on-chip memory and severe workload imbalance; the pre…
Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression
Wei Luo, Yi Huang, Songchen Ma +3
The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when they process long contexts. Curren…