Publications (6)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
Xing Li, Zeyu Xing, Yiming Li +6
KV cache quantization can improve Large Language Models (LLMs) inference throughput and latency in long contexts and large batch-size scenarios while preserving LLMs effectiveness.…
Energy-Efficient Channel Decoding for Wireless Federated Learning: Convergence Analysis and Adaptive Design
Linping Qu, Yuyi Mao, Shenghui Song +1
One of the most critical challenges for deploying distributed learning solutions, such as federated learning (FL), in wireless networks is the limited battery capacity of mobile cl…
FedAQ: Communication-Efficient Federated Edge Learning via Joint Uplink and Downlink Adaptive Quantization
Linping Qu, Shenghui Song, Chi-Ying Tsui
Federated learning (FL) is a powerful machine learning paradigm which leverages the data as well as the computational resources of clients, while protecting clients' data privacy.…
FedDQ: Communication-Efficient Federated Learning with Descending Quantization
Linping Qu, Shenghui Song, Chi-Ying Tsui
Federated learning (FL) is an emerging learning paradigm without violating users' privacy. However, large model size and frequent model aggregation cause serious communication bott…
FedLAM: Low-latency Wireless Federated Learning via Layer-wise Adaptive Modulation
Linping Qu, Shenghui Song, Chi-Ying Tsui
In wireless federated learning (FL), the clients need to transmit the high-dimensional deep neural network (DNN) parameters through bandwidth-limited channels, which causes the com…
How Robust is Federated Learning to Communication Error? A Comparison Study Between Uplink and Downlink Channels
Linping Qu, Shenghui Song, Chi-Ying Tsui +1
Because of its privacy-preserving capability, federated learning (FL) has attracted significant attention from both academia and industry. However, when being implemented over wire…