1 paper
Xiaowei Chi, Junbo Qi, Rongyu Zhang +3
While current LLM chatbots like GPT-4V bridge the gap between human instructions and visual representations to enable text-image generations, they still lack efficient alignment me…