1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Longhao Chen, Yina Zhao, Qiangjun Xie +1
This article optimizes the inference performance of the Qwen-1.8B model by performing Int8 quantization, vectorizing some operators in llama.cpp, and modifying the compilation scri…