#text-to-image generation
12 papers · 1 filter
Amortized Moment Matching for Visual Generation
Wenze Liu, Xintao Wang, Pengfei Wan +1
The paper introduces amortized moment matching, using neural networks to learn data moments as training signals, and proposes the Amortized Fréchet Distance loss to improve one-ste…
Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives
Xiaolong Liu, Junjian Li, Yuan Xiao +4
The paper introduces Dualin, a two‑stage method that simultaneously recovers a human‑readable text prompt and the latent noise of a target image to improve prompt inversion for tex…
Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
Xinyi Wang, Yuyang Huang, Yalin Su +4
The paper introduces AnchorSteer, a training‑free method that improves text‑to‑image diffusion models by initializing with CLIP‑aligned latent noise and actively correcting semanti…
ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform
Xiaoxiao Jiang, Suyi Li, Sheng Yao +5
ServerlessT2I breaks down text-to-image generation pipelines into separate model functions that can be independently scheduled on a serverless platform, allowing per-model scaling…
OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation
Yajing Xu, Yarong Lan, Jiaoyan Chen +6
The paper presents OmniPhys, a knowledge-graph-based benchmark for evaluating physical commonsense in text-to-image models, and OmniPrompt, an iterative optimization framework that…
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
Krzysztof Adamkiewicz, Brian Bernhard Moser, Stanislav Frolov +3
The paper evaluates modern text-to-image diffusion models as sources of synthetic training data and finds that, despite higher visual quality, newer models produce less diverse ima…