4 papers
ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression
Wenya Yu, Chao Zhang, Li Wang +2
Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment. However, applying them sequent…
BAQ: Efficient Bit Allocation Quantization for Large Language Models
Chao Zhang, Li Wang, Samson Lasaulce +1
Post-training model quantization is a widely adopted technique for reducing the memory and computational costs of large language models (LLMs). However, most existing methods rely…
Goal-Oriented State Information Compression for Linear Dynamical System Control
Li Wang, Chao Zhang, Samson Lasaulce +2
In this paper, we consider controlled linear dynamical systems in which the controller has only access to a compressed version of the system state. The technical problem we investi…
Generative AI for RF Sensing in IoT systems
Li Wang, Chao Zhang, Qiyang Zhao +5
The development of wireless sensing technologies, using signals such as Wi-Fi, infrared, and RF to gather environmental data, has significantly advanced within Internet of Things (…