1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 1 cited
FoldGPT: Simple and Effective Large Language Model Compression Scheme
Songwei Liu, Chao Zeng, Lianqiang Li +4
The demand for deploying large language models(LLMs) on mobile devices continues to increase, driven by escalating data security concerns and cloud costs. However, network bandwidt…
cs.LG2024
Differentiable Search for Finding Optimal Quantization Strategy
Lianqiang Li, Chenqian Yan, Yefei Chen
To accelerate and compress deep neural networks (DNNs), many network quantization algorithms have been proposed. Although the quantization strategy of any algorithm from the state-…
cs.AI2023
SparseByteNN: A Novel Mobile Inference Acceleration Framework Based on Fine-Grained Group Sparsity
Haitao Xu, Songwei Liu, Yuyang Xu +7
To address the challenge of increasing network size, researchers have developed sparse models through network pruning. However, maintaining model accuracy while achieving significa…