2 papers
cs.LG2026
CoQuant: Joint Weight-Activation Subspace Projection for Mixed-Precision LLMs
Zhe Ding, Su Pan, Duowei Pan
Post-training quantization (PTQ) has become an important technique for reducing the inference cost of Large Language Models (LLMs). While recent mixed-precision methods improve ult…
cs.CL2026
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
Xinsen Zhang, Zhenkai Ding, Tianjun Pan +4
Extending the effective context length of large language models (LLMs) remains a central challenge for real-world applications. While recent post-training methods have made progres…