54 citations · 180 across the 20 of their papers we have counts for
3 papers · 1 filter
Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement
Zhiyuan Chang, Mingyang Li, Junjie Wang +3
Text-to-Image Diffusion Models (T2I DMs) have garnered significant attention for their ability to generate high-quality images from textual descriptions. However, these models ofte…
VEglue: Testing Visual Entailment Systems via Object-Aligned Joint Erasing
Zhiyuan Chang, Mingyang Li, Junjie Wang +2
Visual entailment (VE) is a multimodal reasoning task consisting of image-sentence pairs whereby a promise is defined by an image, and a hypothesis is described by a sentence. The…
Adversarial Testing for Visual Grounding via Image-Aware Property Reduction
Zhiyuan Chang, Mingyang Li, Junjie Wang +4
Due to the advantages of fusing information from various modalities, multimodal learning is gaining increasing attention. Being a fundamental task of multimodal learning, Visual Gr…