19 citations · 20 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 1 cited
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
Haoran Lin, Xianzhi Yu, Kang Zhao +17
FlashAttention series has been widely applied in the inference of large language models (LLMs). However, FlashAttention series only supports the high-level GPU architectures, e.g.,…
cs.CV2022
CSMPQ:Class Separability Based Mixed-Precision Quantization
Mingkai Wang, Taisong Jin, Miaohui Zhang +1
Mixed-precision quantization has received increasing attention for its capability of reducing the computational burden and speeding up the inference time. Existing methods usually…
cs.CV2022★ 19 cited
TRT-ViT: TensorRT-oriented Vision Transformer
Xin Xia, Jiashi Li, Jie Wu +4
We revisit the existing excellent Transformers from the perspective of practical application. Most of them are not even as efficient as the basic ResNets series and deviate from th…