11 papers
Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment
Huayang Huang, Ruoyu Wang, Jinhui Zhao +5
Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models while preserving perceptual qua…
GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation
Ye Zhu, Kaleb S. Newman, Johannes F. Lutzeyer +3
Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I dive…
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
Yuhan Pei, Ruoyu Wang, Yongqi Yang +3
Originating from the diffusion phenomenon in physics, which describes the random movement and collisions of particles, diffusion generative models simulate a random walk in the dat…
BNMusic: Blending Environmental Noises into Personalized Music
Chi Zuo, Martin B. Møller, Pablo MartÃnez-Nuevo +3
While being disturbed by environmental noises, the acoustic masking technique is a conventional way to reduce the annoyance in audio engineering that seeks to cover up the noises w…
ARC Is a Vision Problem!
Keya Hu, Ali Cy, Linlu Qiu +5
The Abstraction and Reasoning Corpus (ARC) is designed to promote research on abstract reasoning, a fundamental aspect of human intelligence. Common approaches to ARC treat it as a…
D2D: Detector-to-Differentiable Critic for Improved Numeracy in Text-to-Image Generation
Nobline Yoo, Olga Russakovsky, Ye Zhu
Text-to-image (T2I) diffusion models have achieved strong performance in semantic alignment, yet they still struggle with generating the correct number of objects specified in prom…