5 papers · 1 filter
Structure-Guided Visual Perturbation Neutralization for LVLMs
Yuanhe Zhang, Xueting Wang, YanBin Ren +6
Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface through which adversarial pert…
Reward Incremental Learning in Text-to-Image Generation
Maorong Wang, Jiafeng Mao, Xueting Wang +1
The recent success of denoising diffusion models has significantly advanced text-to-image generation. While these large-scale pretrained models show excellent performance in genera…
Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data
Zerun Wang, Jiafeng Mao, Xueting Wang +1
Generative models have become a powerful tool for synthesizing training data in computer vision tasks. Current approaches solely focus on aligning generated images with the target…
The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
Jiafeng Mao, Xueting Wang, Kiyoharu Aizawa
Text-to-image diffusion models allow users control over the content of generated images. Still, text-to-image generation occasionally leads to generation failure requiring users to…
Training-Free Location-Aware Text-to-Image Synthesis
Jiafeng Mao, Xueting Wang
Current large-scale generative models have impressive efficiency in generating high-quality images based on text prompts. However, they lack the ability to precisely control the si…