13 citations · 13 across the 2 of their papers we have counts for
3 papers
cs.AR2024
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
Chenguang Zhang, Zhihang Yuan, Xingchen Li +1
Deep neural networks are widely deployed in many fields. Due to the in-situ computation (known as processing in memory) capacity of the Resistive Random Access Memory (ReRAM) cross…
cs.CL2024★ 13 cited
LLM Inference Unveiled: Survey and Roofline Model Insights
Zhihang Yuan, Yuzhang Shang, Yang Zhou +11
The field of efficient Large Language Model (LLM) inference is rapidly evolving, presenting a unique blend of opportunities and challenges. Although the field has expanded and is v…
cs.CL2023
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
Zhihang Yuan, Yuzhang Shang, Yue Song +4
In this paper, we introduce a new post-training compression paradigm for Large Language Models (LLMs) to facilitate their wider adoption. We delve into LLM weight low-rank decompos…