3 papers
cs.CV2026
FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation
Xueji Fang, Liyuan Ma, Jianhao Zeng +3
Diffusion transformer (DiT) has been widely adopted in the generative diffusion field, advancing the denoising of query tokens through attention and Feed-Forward (\text{FFN}) layer…
cs.CV2026
MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching
Jiahui Huang, Yasi Zhang, Tianyu Chen +6
Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality…
cs.CV2025
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO
Xueji Fang, Liyuan Ma, Zhiyang Chen +2
Recent advances in text-to-video generation, particularly with autoregressive models, have enabled the synthesis of high-quality videos depicting individual scenes. However, extend…