2 papers
cs.CV2025
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
Shijie Zhou, Ruiyi Zhang, Huaisheng Zhu +5
We introduce LLaVA-Reward, an efficient reward model designed to automatically evaluate text-to-image (T2I) generations across multiple perspectives, leveraging pretrained multimod…
cs.CV2024
Enhancing Diffusion Posterior Sampling for Inverse Problems by Integrating Crafted Measurements
Shijie Zhou, Huaisheng Zhu, Rohan Sharma +4
Diffusion models have emerged as a powerful foundation model for visual generations. With an appropriate sampling process, it can effectively serve as a generative prior for solvin…