activity
20212023
most citedOliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization

150 citations · 155 across the 7 of their papers we have counts for

collaborators

7 papers

cs.AR20231 cited

Accelerating Generic Graph Neural Networks via Architecture, Compiler, Partition Method Co-Design

Shuwen Lu, Zhihui Zhang, Cong Guo +3

Graph neural networks (GNNs) have shown significant accuracy improvements in a variety of graph learning domains, sparking considerable research interest. To translate these accura…

cs.DC2023

AdaptGear: Accelerating GNN Training via Adaptive Subgraph-Level Kernels on GPUs

Yangjie Zhou, Yaoxu Song, Jingwen Leng +7

Graph neural networks (GNNs) are powerful tools for exploring and learning from graph structures and features. As such, achieving high-performance execution for GNNs becomes crucia…

cs.AR2023150 cited

OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization

Cong Guo, Jiaming Tang, Weiming Hu +6

Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by every two years, which outpaces the hardware…

cs.AR2023

ImaGen: A General Framework for Generating Memory- and Power-Efficient Image Processing Accelerators

Nisarg Ujjainkar, Jingwen Leng, Yuhao Zhu

Image processing algorithms are prime targets for hardware acceleration as they are commonly used in resource- and power-limited applications. Today's image processing accelerator…

cs.LG20223 cited

ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization

Cong Guo, Chen Zhang, Jingwen Leng +5

Quantization is a technique to reduce the computation and memory cost of DNN models, which are getting increasingly large. Existing quantization solutions use fixed-point integer o…

cs.AR20221 cited

SALO: An Efficient Spatial Accelerator Enabling Hybrid Sparse Attention Mechanisms for Long Sequences

Guan Shen, Jieru Zhao, Quan Chen +3

The attention mechanisms of transformers effectively extract pertinent information from the input sequence. However, the quadratic complexity of self-attention w.r.t the sequence l…