Enhance Image-to-Image Generation with LLaVA-generated Prompts
arXiv:2406.01956 · doi:10.1109/ISPDS62779.2024.10667513
Abstract
This paper presents a novel approach to enhance image-to-image generation by leveraging the multimodal capabilities of the Large Language and Vision Assistant (LLaVA). We propose a framework where LLaVA analyzes input images and generates textual descriptions, hereinafter LLaVA-generated prompts. These prompts, along with the original image, are fed into the image-to-image generation pipeline. This enriched representation guides the generation process towards outputs that exhibit a stronger resemblance to the input image. Extensive experiments demonstrate the effectiveness of LLaVA-generated prompts in promoting image similarity. We observe a significant improvement in the visual coherence between the generated and input images compared to traditional methods. Future work will explore fine-tuning LLaVA prompts for increased control over the creative process. By providing more specific details within the prompts, we aim to achieve a delicate balance between faithfulness to the original image and artistic expression in the generated outputs.
Accepted by 2024 5th International Conference on Information Science, Parallel and Distributed Systems
References in corpus (13)
- Zero-Shot Text-to-Image Generation
- Improved Baselines with Visual Instruction Tuning
- Deep Learning for Content-based Personalized Viewport Prediction of 360-Degree VR Videos
- Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
- AD-Aligning: Emulating Human-like Generalization for Cognitive Domain Adaptation in Deep Learning
- Research on Detection of Floating Objects in River and Lake Based on AI Intelligent Image Recognition
- Research on Image Recognition Technology Based on Multimodal Deep Learning
- CoRMF: Criticality-Ordered Recurrent Mean Field Ising Solver
- WorDepth: Variational Language Prior for Monocular Depth Estimation
- Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models
- Few-shot Name Entity Recognition on StackOverflow
- A Stochastic GDA Method With Backtracking For Solving Nonconvex Concave Minimax Problems
- T-GAE: Transferable Graph Autoencoder for Network Alignment