4 citations · 5 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023
MChat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
Xiaowei Chi, Junbo Qi, Rongyu Zhang +3
While current LLM chatbots like GPT-4V bridge the gap between human instructions and visual representations to enable text-image generations, they still lack efficient alignment me…
cs.CV2023★ 1 cited
Improving Compositional Text-to-image Generation with Large Vision-Language Models
Song Wen, Guian Fang, Renrui Zhang +3
Recent advancements in text-to-image models, particularly diffusion models, have shown significant promise. However, compositional text-to-image models frequently encounter difficu…