most citedTaming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

21 citations · 32 across the 5 of their papers we have counts for

collaborators

5 papers

cs.RO20244 cited

RoboDreamer: Learning Compositional World Models for Robot Imagination

Siyuan Zhou, Yilun Du, Jiaben Chen +3

Text-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environme…

cs.CV20241 cited

Instruct-Imagen: Image Generation with Multi-modal Instruction

Hexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su +9

This paper presents instruct-imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce *multi-modal instruction* for image…

cs.CV20235 cited

Identity Encoder for Personalized Diffusion

Yu-Chuan Su, Kelvin C. K. Chan, Yandong Li +5

Many applications can benefit from personalized image generation models, including image enhancement, video conferences, just to name a few. Existing works achieved personalization…

cs.CV202321 cited

Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

Xuhui Jia, Yang Zhao, Kelvin C. K. Chan +6

This paper proposes a method for generating images of customized objects specified by users. The method is based on a general framework that bypasses the lengthy optimization requi…

cs.LG20231 cited

Learning to Adapt to Online Streams with Distribution Shifts

Chenyan Wu, Yimu Pan, Yandong Li +1

Test-time adaptation (TTA) is a technique used to reduce distribution gaps between the training and testing sets by leveraging unlabeled test data during inference. In this work, w…