21 citations · 24 across the 2 of their papers we have counts for
3 papers
cs.CV2022★ 21 cited
UniCLIP: Unified Framework for Contrastive Language-Image Pre-training
Janghyeon Lee, Jongsuk Kim, Hyounguk Shon +4
Pre-training vision-language models with contrastive objectives has shown promising results that are both scalable to large uncurated datasets and transferable to many downstream a…
cs.CV2022★ 3 cited
MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction Detection
Bumsoo Kim, Jonghwan Mun, Kyoung-Woon On +3
Human-Object Interaction (HOI) detection is the task of identifying a set of <human, object, interaction> triplets from an image. Recent work proposed transformer encoder-decoder a…
cs.CV2021
HOTR: End-to-End Human-Object Interaction Detection with Transformers
Bumsoo Kim, Junhyun Lee, Jaewoo Kang +2
Human-Object Interaction (HOI) detection is a task of identifying "a set of interactions" in an image, which involves the i) localization of the subject (i.e., humans) and target (…