2 papers
cs.DC2026
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
Adrian P. Dieguez, Victor Conchello Vendrell, Alex Batlle +3
Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL library implements it. Yet, TRL treats b…
cs.LG2025
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
Jung Hyun Lee, Seungjae Shin, Vinnam Kim +2
As the rapid scaling of large language models (LLMs) poses significant challenges for deployment on resource-constrained devices, there is growing interest in extremely low-bit qua…