collaborators

11 papers

cs.CV2026

Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment

Huayang Huang, Ruoyu Wang, Jinhui Zhao +5

Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models while preserving perceptual qua…

cs.CV2026

GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

Ye Zhu, Kaleb S. Newman, Johannes F. Lutzeyer +3

Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I dive…

cs.CV2026

SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation

Yuhan Pei, Ruoyu Wang, Yongqi Yang +3

Originating from the diffusion phenomenon in physics, which describes the random movement and collisions of particles, diffusion generative models simulate a random walk in the dat…

cs.SD2026

BNMusic: Blending Environmental Noises into Personalized Music

Chi Zuo, Martin B. Møller, Pablo Martínez-Nuevo +3

While being disturbed by environmental noises, the acoustic masking technique is a conventional way to reduce the annoyance in audio engineering that seeks to cover up the noises w…

cs.CV2025

ARC Is a Vision Problem!

Keya Hu, Ali Cy, Linlu Qiu +5

The Abstraction and Reasoning Corpus (ARC) is designed to promote research on abstract reasoning, a fundamental aspect of human intelligence. Common approaches to ARC treat it as a…

cs.CV2025

D2D: Detector-to-Differentiable Critic for Improved Numeracy in Text-to-Image Generation

Nobline Yoo, Olga Russakovsky, Ye Zhu

Text-to-image (T2I) diffusion models have achieved strong performance in semantic alignment, yet they still struggle with generating the correct number of objects specified in prom…