7 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.CV2024
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
Shihao Zhao, Shaozhe Hao, Bojia Zi +2
Text-to-image generation has made significant advancements with the introduction of text-to-image diffusion models. These models typically consist of a language model that interpre…
cs.CV2023★ 7 cited
MP-Former: Mask-Piloted Transformer for Image Segmentation
Hao Zhang, Feng Li, Huaizhe Xu +4
We present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers fr…