133 citations · 530 across the 144 of their papers we have counts for
7 papers · 1 filter
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
Xin Luo, Jiahao Wang, Chenyuan Wu +6
Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce…
OmniGen2: Towards Instruction-Aligned Multimodal Generation
Chenyuan Wu, Pengfei Zheng, Ruiran Yan +19
In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, imag…
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Junjie Zhou, Zheng Liu, Ze Liu +6
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs,…
NoiseDiffusion: Correcting Noise for Image Interpolation with Diffusion Models beyond Spherical Linear Interpolation
PengFei Zheng, Yonggang Zhang, Zhen Fang +3
Image interpolation based on diffusion models is promising in creating fresh and interesting images. Advanced interpolation methods mainly focus on spherical linear interpolation,…
Invariant Representation via Decoupling Style and Spurious Features from Images
Ruimeng Li, Yuanhao Pu, Zhaoyi Li +2
This paper considers the out-of-distribution (OOD) generalization problem under the setting that both style distribution shift and spurious features exist and domain labels are mis…
Self-Supervised Text Erasing with Controllable Image Synthesis
Gangwei Jiang, Shiyao Wang, Tiezheng Ge +3
Recent efforts on scene text erasing have shown promising results. However, existing methods require rich yet costly label annotations to obtain robust models, which limits the use…