most citedRobust Camera Pose Refinement for Multi-Resolution Hash Encoding

5 citations · 6 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

Bridging Vision and Language Spaces with Assignment Prediction

Jungin Park, Jiyoung Lee, Kwanghoon Sohn

This paper introduces VLAP, a novel approach that bridges pretrained vision models and large language models (LLMs) to make frozen LLMs understand the visual world. VLAP transforms…

cs.CV2023

Dense Text-to-Image Generation with Attention Modulation

Yunji Kim, Jiyoung Lee, Jin-Hwa Kim +2

Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image re…

cs.CV2023

Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning

Hanjae Kim, Jiyoung Lee, Seongheon Park +1

Compositional zero-shot learning (CZSL) aims to recognize unseen compositions with prior knowledge of known primitives (attribute and object). Previous works for CZSL often suffer…

cs.CV2023

Panoramic Image-to-Image Translation

Soohyun Kim, Junho Kim, Taekyung Kim +4

In this paper, we tackle the challenging task of Panoramic Image-to-Image translation (Pano-I2I) for the first time. This task is difficult due to the geometric distortion of panor…

cs.CV2023

Three Recipes for Better 3D Pseudo-GTs of 3D Human Mesh Estimation in the Wild

Gyeongsik Moon, Hongsuk Choi, Sanghyuk Chun +2

Recovering 3D human mesh in the wild is greatly challenging as in-the-wild (ITW) datasets provide only 2D pose ground truths (GTs). Recently, 3D pseudo-GTs have been widely used to…

cs.CV2023

Dual-path Adaptation from Image to Video Transformers

Jungin Park, Jiyoung Lee, Kwanghoon Sohn

In this paper, we efficiently transfer the surpassing representation power of the vision foundation models, such as ViT and Swin, for video understanding with only a few trainable…