collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation

Sangmin Jung, Utkarsh Nath, Yezhou Yang +5

Text-to-image generation models have achieved remarkable capabilities in synthesizing images, but often struggle to provide fine-grained control over the output. Existing guidance…

cs.CV2025

Deep Geometric Moments Promote Shape Consistency in Text-to-3D Generation

Utkarsh Nath, Rajeev Goel, Eun Som Jeon +5

To address the data scarcity associated with 3D assets, 2D-lifting techniques such as Score Distillation Sampling (SDS) have become a widely adopted practice in text-to-3D generati…

cs.CV2024

Steering Rectified Flow Models in the Vector Field for Controlled Image Generation

Maitreya Patel, Song Wen, Dimitris N. Metaxas +1

Diffusion models (DMs) excel in photorealism, image editing, and solving inverse problems, aided by classifier-free guidance and image inversion techniques. However, rectified flow…

cs.CV2024

Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model

Sheng Cheng, Maitreya Patel, Yezhou Yang

Despite advancements in text-to-image models, generating images that precisely align with textual descriptions remains challenging due to misalignment in training data. In this pap…

cs.CV2024

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

Maitreya Patel, Abhiram Kusumba, Sheng Cheng +4

Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the train…

cs.CV2024

R.A.C.E.: Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model

Changhoon Kim, Kyle Min, Yezhou Yang

In the evolving landscape of text-to-image (T2I) diffusion models, the remarkable capability to generate high-quality images from textual descriptions faces challenges with the pot…