1 paper · 1 filter
Eric Tillmann Bill, Enis Simsar, Alessio Tonioni +1
Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified vision--language models derived…