2 papers
cs.LG2025
MSQ: Memory-Efficient Bit Sparsification Quantization
Seokho Han, Seoyeon Yoon, Jinhee Kim +4
As deep neural networks (DNNs) see increased deployment on mobile and edge devices, optimizing model efficiency has become crucial. Mixed-precision quantization is widely favored,…
cs.LG2025
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
Jinhee Kim, Seoyeon Yoon, Taeho Lee +3
The deployment of deep neural networks on edge devices is a challenging task due to the increasing complexity of state-of-the-art models, requiring efforts to reduce model size and…