5 papers
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Xuan He, Cong Wei, Yuhao Cheng +19
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: b…
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Yifu Luo, Haoyuan Sun, Xinhao Hu +12
Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is h…
Watch Before You Answer: Learning from Visually Grounded Post-Training
Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang +8
It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video und…
Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing
Runze He, Yiji Cheng, Tiankai Hang +11
In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful executi…
Group Relative Attention Guidance for Image Editing
Xuanpu Zhang, Xuesong Niu, Ruidong Chen +6
Recently, image editing based on Diffusion-in-Transformer models has undergone rapid development. However, existing editing methods often lack effective control over the degree of…