194 citations · 221 across the 4 of their papers we have counts for
4 papers
A Speed Odyssey for Deployable Quantization of LLMs
Qingyuan Li, Ran Meng, Yiduo Li +6
The large language model era urges faster and less costly inference. Prior model compression works on LLMs tend to undertake a software-centric approach primarily focused on the si…
FPTQ: Fine-grained Post-Training Quantization for Large Language Models
Qingyuan Li, Yifan Zhang, Liang Li +6
In the era of large-scale language models, the substantial parameter size poses significant challenges for deployment. Being a prevalent compression technique, quantization has eme…
EfficientRep:An Efficient Repvgg-style ConvNets with Hardware-aware Neural Network Design
Kaiheng Weng, Xiangxiang Chu, Xiaoming Xu +2
We present a hardware-efficient architecture of convolutional neural network, which has a repvgg-like architecture. Flops or parameters are traditional metrics to evaluate the effi…
YOLOv6 v3.0: A Full-Scale Reloading
Chuyi Li, Lulu Li, Yifei Geng +6
The YOLO community has been in high spirits since our first two releases! By the advent of Chinese New Year 2023, which sees the Year of the Rabbit, we refurnish YOLOv6 with numero…