2 papers
cs.LG2026
REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
Qian Zhang, Yaoming Li, Zhewen Tan +9
Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a…
cs.NE2026
Model Merging to Evolution: Parameter Space Exploration for Expert Models
Chao Wang, Yuchen Guo, Zheng Tan +4
Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training, thereby reducing computational resource…