1 paper
Weiyao Huang, Liqin Wang, Ziqi Sheng +1
Modern image editors combine vision-language models (VLMs) with diffusion transformer backbones to modify a single reference image according to instructions without fine-tuning. Th…