From the 1 of 12 linked papers with an AI index.
12 papers
PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models
Mingde Yao, Zhiyuan You, King-Man Tam +2
PhotoAgent is an autonomous photo‑editing system that interprets a user’s aesthetic intent, plans a sequence of editing actions with tree search, and iteratively refines images usi…
How far have we gone in Generative Image Restoration? A study on its capability, limitations and evaluation practices
Xiang Yin, Jinfan Hu, Zhiyuan You +4
Generative Image Restoration (GIR) has achieved impressive perceptual realism, but how far have its practical capabilities truly advanced compared with previous methods? To answer…
PhotoFramer: Multi-modal Image Composition Instruction
Zhiyuan You, Ke Wang, He Zhang +5
Composition matters during the photo-taking process, yet many casual users struggle to frame well-composed images. To provide composition guidance, we introduce PhotoFramer, a mult…
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
Xin Cai, Zhiyuan You, Zhoutong Zhang +1
Reducing token count is crucial for efficient training and inference of latent diffusion models, especially at high resolution. A common strategy is to build high-compression image…
Position: Evaluation of Visual Processing Should Be Human-Centered, Not Metric-Centered
Jinfan Hu, Fanghua Yu, Zhiyuan You +5
This position paper argues that the evaluation of modern visual processing systems should no longer be driven primarily by single-metric image quality assessment benchmarks, partic…
Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining
Jinfan Hu, Zhiyuan You, Jinjin Gu +3
Generalization to unseen degradations remains a fundamental challenge for low-level vision models. This paper aims to investigate the underlying mechanism of this failure, using im…