5 citations · 9 across the 4 of their papers we have counts for
4 papers
Group Generalized Mean Pooling for Vision Transformer
Byungsoo Ko, Han-Gyu Kim, Byeongho Heo +4
Vision Transformer (ViT) extracts the final representation from either class token or an average of all patch tokens, following the architecture of Transformer in Natural Language…
Back from the future: bidirectional CTC decoding using future information in speech recognition
Namkyu Jung, Geonmin Kim, Han-Gyu Kim
In this paper, we propose a simple but effective method to decode the output of Connectionist Temporal Classifier (CTC) model using a bi-directional neural language model. The bidi…
Proxy Synthesis: Learning with Synthetic Classes for Deep Metric Learning
Geonmo Gu, Byungsoo Ko, Han-Gyu Kim
One of the main purposes of deep metric learning is to construct an embedding space that has well-generalized embeddings on both seen (training) classes and unseen (test) classes.…
Learning with Memory-based Virtual Classes for Deep Metric Learning
Byungsoo Ko, Geonmo Gu, Han-Gyu Kim
The core of deep metric learning (DML) involves learning visual similarities in high-dimensional embedding space. One of the main challenges is to generalize from seen classes of t…