2 papers
cs.LG2026
Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs
Yu Luo, Bo Dong, Wenhua Cheng +1
Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional qua…
cs.LG2023
An Efficient Sparse Inference Software Accelerator for Transformer-based Language Models on CPUs
Haihao Shen, Hengyu Meng, Bo Dong +9
In recent years, Transformer-based language models have become the standard approach for natural language processing tasks. However, stringent throughput and latency requirements i…