3 papers
cs.CV2025
Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing
Taihang Hu, Linxuan Li, Kai Wang +3
Text-to-image generation has seen groundbreaking advancements with diffusion models, enabling high-fidelity synthesis and precise image editing through cross-attention manipulation…
cs.CV2024
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
Taihang Hu, Linxuan Li, Joost van de Weijer +6
Although text-to-image (T2I) models exhibit remarkable generation capabilities, they frequently fail to accurately bind semantically related objects or attributes in the input prom…
cs.CV2024
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
Senmao Li, Taihang Hu, Joost van de Weijer +7
One of the main drawback of diffusion models is the slow inference time for image generation. Among the most successful approaches to addressing this problem are distillation metho…