8 papers
Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition
Pengyang Ling, Jiazi Bu, Yujie Zhou +6
Group Relative Policy Optimization(GRPO) has emerged as an effective paradigm for aligning flow-based generative models with human preferences. However, the high cost of group roll…
MERIT: Multi-domain Efficient RAW Image Translation
Wenjun Huang, Shenghao Fu, Yian Jin +10
RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their d…
Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
Zhixiang Wei, Xiaoxiao Ma, Ruishen Yan +5
Vision Foundation Models(VFMs) have achieved remarkable success in various computer vision tasks. However, their application to semantic segmentation is hindered by two significant…
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
Zhixiang Wei, Guangting Wang, Xiaoxiao Ma +4
Large-scale but noisy image-text pair data have paved the way for the success of Contrastive Language-Image Pretraining (CLIP). As the foundation vision encoder, CLIP in turn serve…
STAR: Scale-wise Text-conditioned AutoRegressive image generation
Xiaoxiao Ma, Mohan Zhou, Tao Liang +5
We introduce STAR, a text-to-image model that employs a scale-wise auto-regressive paradigm. Unlike VAR, which is constrained to class-conditioned synthesis for images up to 256$\t…
Integral Fast Fourier Color Constancy
Wenjun Wei, Yanlin Qian, Huaian Chen +2
Traditional auto white balance (AWB) algorithms typically assume a single global illuminant source, which leads to color distortions in multi-illuminant scenes. While recent neural…