7 papers · 1 filter
MemoGen: Can Past Experience Improve Future Text-to-Image Generation?
Wenshuo Chen, Kuimou Yu, Bowen Tian +10
Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational reasoning, or external knowled…
Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models
Haosen Li, Wenshuo Chen, Lei Wang +4
Diffusion models have achieved remarkable success in synthesizing complex static and temporal visuals, a breakthrough largely driven by Classifier-Free Guidance (CFG). However, des…
Oracle Noise: Faster Semantic Spherical Alignment for Interpretable Latent Optimization
Haosen Li, Wenshuo Chen, Lei Wang +3
Text-to-image diffusion models have achieved remarkable generative capabilities, yet accurately aligning complex textual prompts with synthesized layouts remains an ongoing challen…
-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models
Haosen Li, Wenshuo Chen, Shaofeng Liang +3
Diffusion models have achieved unprecedented success in text-aligned generation, largely driven by Classifier-Free Guidance (CFG). However, standard CFG operates strictly on instan…
Guided Path Sampling: Steering Diffusion Models Back on Track with Principled Path Guidance
Haosen Li, Wenshuo Chen, Shaofeng Liang +3
Iterative refinement methods based on a denoising-inversion cycle are powerful tools for enhancing the quality and control of diffusion models. However, their effectiveness is crit…
POLARIS: Projection-Orthogonal Least Squares for Robust and Adaptive Inversion in Diffusion Models
Wenshuo Chen, Haosen Li, Shaofeng Liang +6
The Inversion-Denoising Paradigm, which is based on diffusion models, excels in diverse image editing and restoration tasks. We revisit its mechanism and reveal a critical, overloo…