4 papers
Human-Corrected Labels Learning: Enhancing Labels Quality via Human Correction of VLMs Discrepancies
Zhongnian Li, Lan Chen, Yixin Xu +2
Vision-Language Models (VLMs), with their powerful content generation capabilities, have been successfully applied to data annotation processes. However, the VLM-generated labels e…
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
Lan Chen, Yuchao Gu, Qi Mao
Large language models, trained on extensive corpora, successfully unify diverse linguistic tasks within a single generative framework. Inspired by this, recent works like Large Vis…
Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
Qi Mao, Lan Chen, Yuchao Gu +2
Balancing fidelity and editability is essential in text-based image editing (TIE), where failures commonly lead to over- or under-editing issues. Existing methods typically rely on…
Edit Transfer: Learning Image Editing via Vision In-Context Relations
Lan Chen, Qi Mao, Yuchao Gu +1
We introduce a new setting, Edit Transfer, where a model learns a transformation from just a single source-target example and applies it to a new query image. While text-based meth…