1 citations · 4 across the 8 of their papers we have counts for
6 papers · 1 filter
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Wonkyun Kim, Changin Choi, Wonseok Lee +1
Stimulated by the sophisticated reasoning capabilities of recent Large Language Models (LLMs), a variety of strategies for bridging video modality have been devised. A prominent st…
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
Jimyeong Kim, Jungwon Park, Wonjong Rhee
In text-to-image personalization, a timely and crucial challenge is the tendency of generated images overfitting to the biases present in the reference images. We initiate our stud…
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
Yeji Song, Jimyeong Kim, Wonhark Park +3
In a surge of text-to-image (T2I) models and their customization methods that generate new images of a user-provided subject, current works focus on alleviating the costs incurred…
On-Off Pattern Encoding and Path-Count Encoding as Deep Neural Network Representations
Euna Jung, Jaekeol Choi, EungGu Yun +1
Understanding the encoded representation of Deep Neural Networks (DNNs) has been a fundamental yet challenging objective. In this work, we focus on two possible directions for anal…
Enhancing Contrastive Learning with Efficient Combinatorial Positive Pairing
Jaeill Kim, Duhun Hwang, Eunjung Lee +3
In the past few years, contrastive learning has played a central role for the success of visual unsupervised representation learning. Around the same time, high-performance non-con…
VNE: An Effective Method for Improving Deep Representation by Manipulating Eigenvalue Distribution
Jaeill Kim, Suhyun Kang, Duhun Hwang +2
Since the introduction of deep learning, a wide scope of representation properties, such as decorrelation, whitening, disentanglement, rank, isotropy, and mutual information, have…