2 papers
cs.DC2026
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
Yaozheng Zhang, Wei Wang, Jie Kong +5
The increasing adoption of large language models (LLMs) on heterogeneous computing platforms poses significant challenges to achieving high inference efficiency. To address these e…
cs.DC2025
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
Jie Kong, Junxiang Zhang, Jiheng Xu +7
In the field of deep learning, traditional attention mechanisms face significant challenges related to high computational complexity and large memory consumption when processing lo…