activity
20242026
collaborators

7 papers

cs.CV2026

MiniWorld: Democratizing the Training of Video World Models from Scratch

Yian Zhao, Ruochong Zheng, Hongcan Guo +3

Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions…

cs.CV2026

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

Zitong Xu, Huiyu Duan, Yifei Nie +9

Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as unnatural objects, lighting…

cs.CV2026

Video Generation with Predictive Latents

Yian Zhao, Feng Wang, Qiushan Guo +4

Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, improving training efficiency an…

cs.CV2026

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing

Ying Zeng, Miaosen Luo, Guangyuan Li +10

Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and ca…

cs.CV2025

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment

Zhuoxuan Cai, Jian Zhang, Xinbin Yuan +7

Recent studies demonstrate that multimodal large language models (MLLMs) can proficiently evaluate visual quality through interpretable assessments. However, existing approaches ty…

cs.AI2025

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

Xinbin Yuan, Jian Zhang, Kaixin Li +8

Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…