4 papers
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
Eric Tillmann Bill, Enis Simsar, Alessio Tonioni +1
Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified vision--language models derived…
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
Eric Tillmann Bill, Enis Simsar, Thomas Hofmann
Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. W…
JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models
Eric Tillmann Bill, Enis Simsar, Thomas Hofmann
We introduce JEDI, a test-time adaptation method that enhances subject separation and compositional alignment in diffusion models without requiring retraining or external supervisi…
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
Eric Tillman Bill, Cristian Perez Jensen, Sotiris Anagnostidis +1
Denoising diffusion models exhibit remarkable generative capabilities, but remain challenging to train due to their inherent stochasticity, where high-variance gradient estimates l…