diffusion models 1image editing 1latent noise inversion 1prompt inversion 1text-to-image generation 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives
Xiaolong Liu, Junjian Li, Yuan Xiao +4
The paper introduces Dualin, a two‑stage method that simultaneously recovers a human‑readable text prompt and the latent noise of a target image to improve prompt inversion for tex…
cs.CV2026
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
Jiaqi Deng, Zonghan Wu, Huan Huo +1
Knowledge-based Vision Question Answering (KB-VQA) extends general Vision Question Answering (VQA) by not only requiring the understanding of visual and textual inputs but also ext…
cs.CV2025
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
Jiaqi Deng, Kaize Shi, Zonghan Wu +3
Knowledge-based Vision Question Answering (KB-VQA) systems address complex visual-grounded questions with knowledge retrieved from external knowledge bases. The tasks of knowledge…