activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2025

GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

Jing Wang, Jiajun Liang, Jie Liu +10

Recently, GRPO-based reinforcement learning has shown remarkable progress in optimizing flow-matching models, effectively improving their alignment with task-specific rewards. With…

cs.CV2025

PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models

Runze He, Bo Cheng, Yuhang Ma +7

In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images. Unlike previous diff…

cs.CV2025

U-StyDiT: Ultra-high Quality Artistic Style Transfer Using Diffusion Transformers

Zhanjie Zhang, Ao Ma, Ke Cao +6

Ultra-high quality artistic style transfer refers to repainting an ultra-high quality content image using the style information learned from the style image. Existing artistic styl…

cs.CV2025

WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation

Jing Wang, Ao Ma, Ke Cao +9

Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle…

cs.CV2025

RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers

Ke Cao, Jing Wang, Ao Ma +11

The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled di…

cs.CV2024

HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation

Bo Cheng, Yuhang Ma, Liebucha Wu +5

The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout…