6 papers
Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics
Xiaoyan Xing, Xiao Zhang, Sezer Karaoglu +2
Image-to-image relighting requires representations that separate illumination from scene properties while preserving dense geometry, material, and photometric cues. We use this tas…
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
Guozhen Zhang, Xuerui Qiu, Yutao Cui +11
Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDR…
VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement
Shulian Zhang, Yong Guo, Long Peng +6
Video Face Enhancement (VFE) aims to restore high-quality facial regions from degraded video sequences, enabling a wide range of practical applications. Despite substantial progres…
An Item is Worth a Prompt: Versatile Image Editing with Disentangled Control
Aosong Feng, Weikang Qiu, Jinbin Bai +5
Building on the success of text-to-image diffusion models (DPMs), image editing is an important application to enable human interaction with AI-generated content. Among various edi…
InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
Nirat Saini, Navaneeth Bodla, Ashish Shrivastava +4
We introduce InVi, an approach for inserting or replacing objects within videos (referred to as inpainting) using off-the-shelf, text-to-image latent diffusion models. InVi targets…
Integrating View Conditions for Image Synthesis
Jinbin Bai, Zhen Dong, Aosong Feng +3
In the field of image processing, applying intricate semantic modifications within existing images remains an enduring challenge. This paper introduces a pioneering framework that…