5 citations · 10 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 2 cited
Large Language Models are Temporal and Causal Reasoners for Video Question Answering
Dohwan Ko, Ji Soo Lee, Wooyoung Kang +2
Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks. We observe that the LLMs provide effective p…
cs.CV2023★ 3 cited
NICE: CVPR 2023 Challenge on Zero-shot Image Captioning
Taehoon Kim, Pyunghwan Ahn, Sangyun Kim +39
In this report, we introduce NICE (New frontiers for zero-shot Image Captioning Evaluation) project and share the results and outcomes of 2023 challenge. This project is designed t…
cs.CV2023★ 5 cited
Open-Vocabulary Object Detection using Pseudo Caption Labels
Han-Cheol Cho, Won Young Jhoo, Wooyoung Kang +1
Recent open-vocabulary detection methods aim to detect novel objects by distilling knowledge from vision-language models (VLMs) trained on a vast amount of image-text pairs. To imp…