11 citations · 35 across the 7 of their papers we have counts for
3 papers · 1 filter
MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction Detection
Bumsoo Kim, Jonghwan Mun, Kyoung-Woon On +3
Human-Object Interaction (HOI) detection is the task of identifying a set of <human, object, interaction> triplets from an image. Recent work proposed transformer encoder-decoder a…
Geodesic Multi-Modal Mixup for Robust Fine-Tuning
Changdae Oh, Junhyuk So, Hoyoon Byun +4
Pre-trained multi-modal models, such as CLIP, provide transferable embeddings and show promising results in diverse applications. However, the analysis of learned multi-modal embed…
Boundary-aware Self-supervised Learning for Video Scene Segmentation
Jonghwan Mun, Minchul Shin, Gunsoo Han +4
Self-supervised learning has drawn attention through its effectiveness in learning in-domain representations with no ground-truth annotations; in particular, it is shown that prope…