activity
20182022
most citedDetecting Human-Object Interactions with Action Co-occurrence Priors

14 citations · 20 across the 3 of their papers we have counts for

collaborators

8 papers

cs.CV2024

Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality

Youngtaek Oh, Jae Won Cho, Dong-Jin Kim +2

In this paper, we propose a new method to enhance compositional understanding in pre-trained vision and language models (VLMs) without sacrificing performance in zero-shot multi-mo…

cs.CV20226 cited

Signing Outside the Studio: Benchmarking Background Robustness for Continuous Sign Language Recognition

Youngjoon Jang, Youngtaek Oh, Jae Won Cho +3

The goal of this work is background-robust continuous sign language recognition. Most existing Continuous Sign Language Recognition (CSLR) benchmarks have fixed backgrounds and are…

cs.CV2021

Dealing with Missing Modalities in the Visual Question Answer-Difference Prediction Task through Knowledge Distillation

Jae Won Cho, Dong-Jin Kim, Jinsoo Choi +2

In this work, we address the issues of missing modalities that have arisen from the Visual Question Answer-Difference prediction task and find a novel method to solve the task at h…

cs.CV2020

Dense Relational Image Captioning via Multi-task Triple-Stream Networks

Dong-Jin Kim, Tae-Hyun Oh, Jinsoo Choi +1

We introduce dense relational captioning, a novel image captioning task which aims to generate multiple captions with respect to relational information between objects in a visual…

cs.CV202014 cited

Detecting Human-Object Interactions with Action Co-occurrence Priors

Dong-Jin Kim, Xiao Sun, Jinsoo Choi +2

A common problem in human-object interaction (HOI) detection task is that numerous HOI classes have only a small number of labeled examples, resulting in training sets with a long-…

cs.CV2019

Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach

Dong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh +1

Constructing an organized dataset comprised of a large number of images and several captions for each image is a laborious task, which requires vast human effort. On the other hand…