activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition

Pengyang Ling, Jiazi Bu, Yujie Zhou +6

Group Relative Policy Optimization(GRPO) has emerged as an effective paradigm for aligning flow-based generative models with human preferences. However, the high cost of group roll…

cs.CV2025

Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models

Zhixiang Wei, Xiaoxiao Ma, Ruishen Yan +5

Vision Foundation Models(VFMs) have achieved remarkable success in various computer vision tasks. However, their application to semantic segmentation is hindered by two significant…

cs.CV2025

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models

Zhixiang Wei, Guangting Wang, Xiaoxiao Ma +4

Large-scale but noisy image-text pair data have paved the way for the success of Contrastive Language-Image Pretraining (CLIP). As the foundation vision encoder, CLIP in turn serve…

cs.CV2025

STAR: Scale-wise Text-conditioned AutoRegressive image generation

Xiaoxiao Ma, Mohan Zhou, Tao Liang +5

We introduce STAR, a text-to-image model that employs a scale-wise auto-regressive paradigm. Unlike VAR, which is constrained to class-conditioned synthesis for images up to 256$\t…

cs.CV2024

Masked Pre-training Enables Universal Zero-shot Denoiser

Xiaoxiao Ma, Zhixiang Wei, Yi Jin +5

In this work, we observe that model trained on vast general images via masking strategy, has been naturally embedded with their distribution knowledge, thus spontaneously attains t…

cs.CV2024

MotionClone: Training-Free Motion Cloning for Controllable Video Generation

Pengyang Ling, Jiazi Bu, Pan Zhang +6

Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular…