7 papers
MiniWorld: Democratizing the Training of Video World Models from Scratch
Yian Zhao, Ruochong Zheng, Hongcan Guo +3
Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions…
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
Zitong Xu, Huiyu Duan, Yifei Nie +9
Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as unnatural objects, lighting…
Video Generation with Predictive Latents
Yian Zhao, Feng Wang, Qiushan Guo +4
Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, improving training efficiency an…
SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
Ying Zeng, Miaosen Luo, Guangyuan Li +10
Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and ca…
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
Zhuoxuan Cai, Jian Zhang, Xinbin Yuan +7
Recent studies demonstrate that multimodal large language models (MLLMs) can proficiently evaluate visual quality through interpretable assessments. However, existing approaches ty…
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
Xinbin Yuan, Jian Zhang, Kaixin Li +8
Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…