9 citations · 24 across the 6 of their papers we have counts for
8 papers
MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction Detection
Bumsoo Kim, Jonghwan Mun, Kyoung-Woon On +3
Human-Object Interaction (HOI) detection is the task of identifying a set of <human, object, interaction> triplets from an image. Recent work proposed transformer encoder-decoder a…
Boundary-aware Self-supervised Learning for Video Scene Segmentation
Jonghwan Mun, Minchul Shin, Gunsoo Han +4
Self-supervised learning has drawn attention through its effectiveness in learning in-domain representations with no ground-truth annotations; in particular, it is shown that prope…
Winning the ICCV'2021 VALUE Challenge: Task-aware Ensemble and Transfer Learning with Visual Concepts
Minchul Shin, Jonghwan Mun, Kyoung-Woon On +3
The VALUE (Video-And-Language Understanding Evaluation) benchmark is newly introduced to evaluate and analyze multi-modal representation learning algorithms on three video-and-lang…
RTIC: Residual Learning for Text and Image Composition using Graph Convolutional Network
Minchul Shin, Yoonjae Cho, Byungsoo Ko +1
In this paper, we study the compositional learning of images and texts for image retrieval. The query is given in the form of an image and text that describes the desired modificat…
Semi-supervised Learning with a Teacher-student Network for Generalized Attribute Prediction
Minchul Shin
This paper presents a study on semi-supervised learning to solve the visual attribute prediction problem. In many applications of vision algorithms, the precise recognition of visu…
Fashion-IQ 2020 Challenge 2nd Place Team's Solution
Minchul Shin, Yoonjae Cho, Seongwuk Hong
This paper is dedicated to team VAA's approach submitted to the Fashion-IQ challenge in CVPR 2020. Given a pair of the image and the text, we present a novel multimodal composition…