1 paper
Longhao Chen, Yina Zhao, Qiangjun Xie +1
This article optimizes the inference performance of the Qwen-1.8B model by performing Int8 quantization, vectorizing some operators in llama.cpp, and modifying the compilation scri…