1 paper
Tianshuo Yuan, Yuxiang Lin, Jue Wang +5
Combining Vision Large Language Models (VLLMs) with diffusion models offers a powerful method for executing image editing tasks based on human language instructions. However, langu…