180 citations · 187 across the 7 of their papers we have counts for
6 papers
Dense Text-to-Image Generation with Attention Modulation
Yunji Kim, Jiyoung Lee, Jin-Hwa Kim +2
Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image re…
Panoramic Image-to-Image Translation
Soohyun Kim, Junho Kim, Taekyung Kim +4
In this paper, we tackle the challenging task of Panoramic Image-to-Image translation (Pano-I2I) for the first time. This task is difficult due to the geometric distortion of panor…
Text-Conditioned Sampling Framework for Text-to-Image Generation with Masked Generative Models
Jaewoong Lee, Sangwon Jang, Jaehyeong Jo +5
Token-based masked generative models are gaining popularity for their fast inference time with parallel decoding. While recent token-based approaches achieve competitive performanc…
Robust Camera Pose Refinement for Multi-Resolution Hash Encoding
Hwan Heo, Taekyung Kim, Jiyoung Lee +4
Multi-resolution hash encoding has recently been proposed to reduce the computational cost of neural renderings, such as NeRF. This method requires accurate camera poses for the ne…
Semi-Parametric Video-Grounded Text Generation
Sungdong Kim, Jin-Hwa Kim, Jiyoung Lee +1
Efficient video-language modeling should consider the computational cost because of a large, sometimes intractable, number of video frames. Parametric approaches such as the attent…
Hadamard Product for Low-rank Bilinear Pooling
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim +3
Bilinear models provide rich representations compared with linear models. They have been applied in various visual tasks, such as object recognition, segmentation, and visual quest…