activity
20202022
most citedS-SGD: Symmetrical Stochastic Gradient Descent with Weight Noise Injection for Reaching Flat Minima

5 citations · 9 across the 6 of their papers we have counts for

collaborators

7 papers

eess.AS20222 cited

A Comparison of Transformer, Convolutional, and Recurrent Neural Networks on Phoneme Recognition

Kyuhong Shim, Wonyong Sung

Phoneme recognition is a very important part of speech recognition that requires the ability to extract phonetic features from multiple frames. In this paper, we compare and analyz…

cs.CL2022

Korean Tokenization for Beam Search Rescoring in Speech Recognition

Kyuhong Shim, Hyewon Bae, Wonyong Sung

The performance of automatic speech recognition (ASR) models can be greatly improved by proper beam-search decoding with external language model (LM). There has been an increasing…

cs.CL20211 cited

Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling

Kyuhong Shim, Iksoo Choi, Wonyong Sung +1

While Transformer-based models have shown impressive language modeling performance, the large computation cost is often prohibitive for practical use. Attention head pruning, which…

cs.LG2020

Stochastic Precision Ensemble: Self-Knowledge Distillation for Quantized Deep Neural Networks

Yoonho Boo, Sungho Shin, Jungwook Choi +1

The quantization of deep neural networks (QDNNs) has been actively studied for deployment in edge devices. Recent studies employ the knowledge distillation (KD) method to improve t…

cs.LG20205 cited

S-SGD: Symmetrical Stochastic Gradient Descent with Weight Noise Injection for Reaching Flat Minima

Wonyong Sung, Iksoo Choi, Jinhwan Park +2

The stochastic gradient descent (SGD) method is most widely used for deep neural network (DNN) training. However, the method does not always converge to a flat minimum of the loss…

cs.LG2020

Quantized Neural Networks: Characterization and Holistic Optimization

Yoonho Boo, Sungho Shin, Wonyong Sung

Quantized deep neural networks (QDNNs) are necessary for low-power, high throughput, and embedded applications. Previous studies mostly focused on developing optimization methods f…