Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression
Wenya Yu, Chao Zhang, Li Wang +2
Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment. However, applying them sequent…
cs.LG2025
BAQ: Efficient Bit Allocation Quantization for Large Language Models
Chao Zhang, Li Wang, Samson Lasaulce +1
Post-training model quantization is a widely adopted technique for reducing the memory and computational costs of large language models (LLMs). However, most existing methods rely…