activity
20182022
most citedLearning Low-Rank Approximation for CNNs

13 citations · 30 across the 5 of their papers we have counts for

collaborators

12 papers

cs.LG20221 cited

AlphaTuning: Quantization-Aware Parameter-Efficient Adaptation of Large-Scale Pre-Trained Language Models

Se Jung Kwon, Jeonghoon Kim, Jeongin Bae +7

There are growing interests in adapting large-scale language models using parameter-efficient fine-tuning methods. However, accelerating the model itself and achieving better infer…

eess.SY20225 cited

DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation

Seongmin Hong, Seungjae Moon, Junsoo Kim +4

Transformer is a deep learning language model widely used for natural language processing (NLP) services in datacenters. Among transformer models, Generative Pre-trained Transforme…

cs.LG2021

Modulating Regularization Frequency for Efficient Compression-Aware Model Training

Dongsoo Lee, Se Jung Kwon, Byeongwook Kim +3

While model compression is increasingly important because of large neural network size, compression-aware training is challenging as it needs sophisticated model modifications and…

cs.LG2021

Q-Rater: Non-Convex Optimization for Post-Training Uniform Quantization

Byeongwook Kim, Dongsoo Lee, Yeonju Ro +4

Various post-training uniform quantization methods have usually been studied based on convex optimization. As a result, most previous ones rely on the quantization error minimizati…

cs.LG2020

FleXOR: Trainable Fractional Quantization

Dongsoo Lee, Se Jung Kwon, Byeongwook Kim +3

Quantization based on the binary codes is gaining attention because each quantized bit can be directly utilized for computations without dequantization using look-up tables. Previo…

cs.LG2020

Extremely Low Bit Transformer Quantization for On-Device Neural Machine Translation

Insoo Chung, Byeongwook Kim, Yoonjung Choi +5

The deployment of widely used Transformer architecture is challenging because of heavy computation load and memory overhead during inference, especially when the target device is l…