200 citations · 374 across the 7 of their papers we have counts for
7 papers
Large Language Models are Temporal and Causal Reasoners for Video Question Answering
Dohwan Ko, Ji Soo Lee, Wooyoung Kang +2
Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks. We observe that the LLMs provide effective p…
CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training
Kihyun You, Jawook Gu, Jiyeon Ham +5
A large-scale image-text pair dataset has greatly contributed to the development of vision-language pre-training (VLP) models, which enable zero-shot or few-shot classification wit…
NICE: CVPR 2023 Challenge on Zero-shot Image Captioning
Taehoon Kim, Pyunghwan Ahn, Sangyun Kim +39
In this report, we introduce NICE (New frontiers for zero-shot Image Captioning Evaluation) project and share the results and outcomes of 2023 challenge. This project is designed t…
Open-Vocabulary Object Detection using Pseudo Caption Labels
Han-Cheol Cho, Won Young Jhoo, Wooyoung Kang +1
Recent open-vocabulary detection methods aim to detect novel objects by distilling knowledge from vision-language models (VLMs) trained on a vast amount of image-text pairs. To imp…
MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models
Dohwan Ko, Joonmyung Choi, Hyeong Kyu Choi +3
Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase,…
PVANet: Lightweight Deep Neural Networks for Real-time Object Detection
Sanghoon Hong, Byungseok Roh, Kye-Hyeon Kim +2
In object detection, reducing computational cost is as important as improving accuracy for most practical usages. This paper proposes a novel network structure, which is an order o…