activity
20192026
most citedProgressive DNN Compression: A Key to Achieve Ultra-High Weight Pruning and Quantization Rates using ADMM

26 citations · 29 across the 6 of their papers we have counts for

collaborators

12 papers

cs.AR2026

Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators

Zongle Huang, Hongyang Jia, Kaiwei Zou +1

Neural network (NN) accelerators with multi-chip-module (MCM) architectures enable integration of massive computation capability; however, they face challenges of computing resourc…

cs.LG2025

SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference

Wenxun Wang, Shuchang Zhou, Wenyu Sun +2

Transformers have shown remarkable performance in both natural language processing (NLP) and computer vision (CV) tasks. However, their real-time inference speed and efficiency are…

cs.LG2025

MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE

Zongle Huang, Lei Zhu, Zongyuan Zhan +5

Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional…

cs.DC2025

Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism

Xinyuan Lin, Chenlu Li, Zongle Huang +5

Larger model sizes and longer sequence lengths have empowered the Large Language Model (LLM) to achieve outstanding performance across various domains. However, this progress bring…

cs.LG2022

Block-Wise Dynamic-Precision Neural Network Training Acceleration via Online Quantization Sensitivity Analytics

Ruoyang Liu, Chenhan Wei, Yixiong Yang +3

Data quantization is an effective method to accelerate neural network training and reduce power consumption. However, it is challenging to perform low-bit quantized training: the c…

cs.ET2021

Enabling Lower-Power Charge-Domain Nonvolatile In-Memory Computing with Ferroelectric FETs

Guodong Yin, Yi Cai, Juejian Wu +6

Compute-in-memory (CiM) is a promising approach to alleviating the memory wall problem for domain-specific applications. Compared to current-domain CiM solutions, charge-domain CiM…