26 citations · 29 across the 6 of their papers we have counts for
12 papers
Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
Zongle Huang, Hongyang Jia, Kaiwei Zou +1
Neural network (NN) accelerators with multi-chip-module (MCM) architectures enable integration of massive computation capability; however, they face challenges of computing resourc…
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
Wenxun Wang, Shuchang Zhou, Wenyu Sun +2
Transformers have shown remarkable performance in both natural language processing (NLP) and computer vision (CV) tasks. However, their real-time inference speed and efficiency are…
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
Zongle Huang, Lei Zhu, Zongyuan Zhan +5
Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional…
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
Xinyuan Lin, Chenlu Li, Zongle Huang +5
Larger model sizes and longer sequence lengths have empowered the Large Language Model (LLM) to achieve outstanding performance across various domains. However, this progress bring…
Block-Wise Dynamic-Precision Neural Network Training Acceleration via Online Quantization Sensitivity Analytics
Ruoyang Liu, Chenhan Wei, Yixiong Yang +3
Data quantization is an effective method to accelerate neural network training and reduce power consumption. However, it is challenging to perform low-bit quantized training: the c…
Enabling Lower-Power Charge-Domain Nonvolatile In-Memory Computing with Ferroelectric FETs
Guodong Yin, Yi Cai, Juejian Wu +6
Compute-in-memory (CiM) is a promising approach to alleviating the memory wall problem for domain-specific applications. Compared to current-domain CiM solutions, charge-domain CiM…