1 citations · 1 across the 3 of their papers we have counts for
3 papers
TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
Junhan Kim, Yeo Jeong Park, Seungwoo Son +4
The rapid growth of large language models (LLMs) has heightened the importance of post-training quantization (PTQ) for reducing memory and computation costs. Among PTQ methods, GPT…
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
Junhan Kim, Chungman Lee, Eulrang Cho +4
With the increasing complexity of generative AI models, post-training quantization (PTQ) has emerged as a promising solution for deploying hyper-scale models on edge devices such a…
Practical Knowledge Distillation: Using DNNs to Beat DNNs
Chung-Wei Lee, Pavlos Athanasios Apostolopulos, Igor L. Markov
For tabular data sets, we explore data and model distillation, as well as data denoising. These techniques improve both gradient-boosting models and a specialized DNN architecture.…