6 citations · 14 across the 6 of their papers we have counts for
6 papers · 1 filter
Revealing Multi-View Hallucination in Large Vision-Language Models
Wooje Park, Insu Lee, Soohyun Kim +4
Large vision-language models (LVLMs) are increasingly being applied to multi-view image inputs captured from diverse viewpoints. Despite this growing use, current LVLMs often gener…
Diffusion-driven GAN Inversion for Multi-Modal Face Image Generation
Jihyun Kim, Changjae Oh, Hoseok Do +2
We present a new multi-modal face image generation method that converts a text prompt and a visual input, such as a semantic mask or scribble map, into a photo-realistic face image…
InstaFormer: Instance-Aware Image-to-Image Translation with Transformer
Soohyun Kim, Jongbeom Baek, Jihye Park +2
We present a novel Transformer-based network architecture for instance-aware image-to-image translation, dubbed InstaFormer, to effectively integrate global- and instance-level inf…
An Attention-based Method for Action Unit Detection at the 3rd ABAW Competition
Duy Le Hoai, Eunchae Lim, Eunbin Choi +5
Facial Action Coding System is an approach for modeling the complexity of human emotional expression. Automatic action unit (AU) detection is a crucial research area in human-compu…
Deep Translation Prior: Test-time Training for Photorealistic Style Transfer
Sunwoo Kim, Soohyun Kim, Seungryong Kim
Recent techniques to solve photorealistic style transfer within deep convolutional neural networks (CNNs) generally require intensive training from large-scale datasets, thus havin…
Online Exemplar Fine-Tuning for Image-to-Image Translation
Taewon Kang, Soohyun Kim, Sunwoo Kim +1
Existing techniques to solve exemplar-based image-to-image translation within deep convolutional neural networks (CNNs) generally require a training phase to optimize the network p…