2 papers
cs.LG2025
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
Jinhao Zhang, Yunquan Zhang, Boyang Zhang +2
Quantization method plays a crucial role in improving model efficiency and reducing deployment costs, enabling the widespread application of deep learning models on resource-constr…
cs.LG2025
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
Zeyu Liu, Yan Li, Yunquan Zhang +6
Training large language models typically demands extensive GPU memory and substantial financial investment, which poses a barrier for many small- to medium-sized teams. In this pap…