papers

Publications (32)

cs.CV2022

NeRF-Art: Text-Driven Neural Radiance Fields Stylization

Can Wang, Ruixiang Jiang, Menglei Chai +3

As a powerful representation of 3D scenes, the neural radiance field (NeRF) enables high-quality novel view synthesis from multi-view images. Stylizing NeRF, however, remains chall…

cs.CV2020

One-Shot Identity-Preserving Portrait Reenactment

Sitao Xiang, Yuming Gu, Pengda Xiang +4

We present a deep learning-based framework for portrait reenactment from a single picture of a target (one-shot) and a video of a driving subject. Existing facial reenactment metho…

cs.CV2026

Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models

Tianyi Xiang, Mingming He, Li Ma +1

Cinematic compositing aims to integrate green-screen characters into novel environments while maintaining physical and photometric realism. Previous methods often fail to capture t…

cs.CV2026

ReAge3D: Re-Aging 3D Faces with View Consistency

Libing Zeng, Li Ma, Mingming He +3

We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity-preserving results. Existing 3D editing methods, while effecti…

cs.GR2022

Water Simulation and Rendering from a Still Photograph

Ryusuke Sugimoto, Mingming He, Jing Liao +1

We propose an approach to simulate and render realistic water animation from a single still input photograph. We first segment the water surface, estimate rendering parameters, and…

cs.GR2025

Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture

Julien Philip, Li Ma, Pascal Clausen +8

We present a unique system for large-scale, multi-performer, high resolution 4D volumetric capture providing realistic free-viewpoint video up to and including 4K resolution facial…

cs.CV2025

Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

Ryan Burgert, Yuancheng Xu, Wenqi Xian +10

Generative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise s…

cs.GR2025

Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset

Yiqun Mei, Mingming He, Li Ma +9

Video portrait relighting remains challenging because the results need to be both photorealistic and temporally stable. This typically requires a strong model design that can captu…

cs.CV2021

Efficient Semantic Image Synthesis via Class-Adaptive Normalization

Zhentao Tan, Dongdong Chen, Qi Chu +6

Spatially-adaptive normalization (SPADE) is remarkably successful recently in conditional semantic image synthesis \cite{park2019semantic}, which modulates the normalized activatio…

cs.CV2026

ID-V2V: Identity-Preserving Video Restylization

Yuancheng Xu, Mingming He, Pablo Salamanca +5

In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual e…

cs.CV2018

Deep Exemplar-based Colorization

Mingming He, Dongdong Chen, Jing Liao +2

We propose the first deep learning approach for exemplar-based local colorization. Given a reference color image, our convolutional neural network directly maps a grayscale image t…

cs.CV2022

DenseGAP: Graph-Structured Dense Correspondence Learning with Anchor Points

Zhengfei Kuang, Jiaman Li, Mingming He +2

Establishing dense correspondence between two images is a fundamental computer vision problem, which is typically tackled by matching local feature descriptors. However, without gl…

cs.LG2021

DisUnknown: Distilling Unknown Factors for Disentanglement Learning

Sitao Xiang, Yuming Gu, Pengda Xiang +4

Disentangling data into interpretable and independent factors is critical for controllable generation tasks. With the availability of labeled data, supervision can help enforce the…

physics.flu-dyn2024

Hydrodynamics of polydisperse gas-solid flows: Kinetic theory and multifluid simulation

Bidan Zhao, Kun Shi, Mingming He +1

Polydisperse gas-solid flows, which is notoriously difficult to model due to the complex gas-particle and particle-particle interactions, are widely encountered in industry. In thi…

cs.CV2024

Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM

Can Wang, Hongliang Zhong, Menglei Chai +3

Automatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), rece…

cs.CV2024

Fitting Spherical Gaussians to Dynamic HDRI Sequences

Pascal Clausen, Li Ma, Mingming He +3

We present a technique for fitting high dynamic range illumination (HDRI) sequences using anisotropic spherical Gaussians (ASGs) while preserving temporal consistency in the compre…

cs.CV2021

Exemplar-Based 3D Portrait Stylization

Fangzhou Han, Shuquan Ye, Mingming He +2

Exemplar-based portrait stylization is widely attractive and highly desired. Despite recent successes, it remains challenging, especially when considering both texture and geometri…

physics.flu-dyn2024

A critical comparison of the implementation of granular pressure gradient term in Euler-Euler simulation of gas-solid flows

Yige Liu, Mingming He, Jianhua Chen +4

Numerical solution of Euler-Euler model using different in-house, open source and commercial software can generate significantly different results, even when the governing equation…

cs.CV2026

BodyReLux: Temporally Consistent Full-Body Video Relighting

Li Ma, Mingming He, Xueming Yu +4

Being able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject-specific video diffusion-based framework for…

cs.CV2019

Deep Exemplar-based Video Colorization

Bo Zhang, Mingming He, Jing Liao +4

This paper presents the first end-to-end network for exemplar-based video colorization. The main challenge is to achieve temporal consistency while remaining faithful to the refere…

cs.CV2026

DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models

Zhengming Yu, Li Ma, Mingming He +11

Most digital videos are stored in 8-bit low dynamic range (LDR) formats, where much of the original high dynamic range (HDR) scene radiance is lost due to saturation and quantizati…

cs.CV2018

Progressive Color Transfer with Dense Semantic Correspondences

Mingming He, Jing Liao, Dongdong Chen +2

We propose a new algorithm for color transfer between images that have perceptually similar semantic structures. We aim to achieve a more accurate color transfer that leverages sem…

cs.CV2024

DifFRelight: Diffusion-Based Facial Performance Relighting

Mingming He, Pascal Clausen, Ahmet Levent Taşel +8

We present a novel framework for free-viewpoint facial performance relighting using diffusion-based image-to-image translation. Leveraging a subject-specific dataset containing div…

cs.CV2022

CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields

Can Wang, Menglei Chai, Mingming He +2

We present CLIP-NeRF, a multi-modal 3D object manipulation method for neural radiance fields (NeRF). By leveraging the joint language-image embedding space of the recent Contrastiv…

cs.CV2022

Cross-Domain and Disentangled Face Manipulation with 3D Guidance

Can Wang, Menglei Chai, Mingming He +2

Face image manipulation via three-dimensional guidance has been widely applied in various interactive scenarios due to its semantically-meaningful understanding and user-friendly c…

cs.CV2020

Rethinking Spatially-Adaptive Normalization

Zhentao Tan, Dongdong Chen, Qi Chu +5

Spatially-adaptive normalization is remarkably successful recently in conditional semantic image synthesis, which modulates the normalized activation with spatially-varying transfo…

cs.GR2020

Dynamic Facial Asset and Rig Generation from a Single Scan

Jiaman Li, Zhengfei Kuang, Yajie Zhao +3

The creation of high-fidelity computer-generated (CG) characters used in film and gaming requires intensive manual labor and a comprehensive set of facial assets to be captured wit…

cs.CV2025

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

Yuancheng Xu, Wenqi Xian, Li Ma +10

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train th…

cs.CV2023

Mesh-Guided Neural Implicit Field Editing

Can Wang, Mingming He, Menglei Chai +2

Neural implicit fields have emerged as a powerful 3D representation for reconstructing and rendering photo-realistic views, yet they possess limited editability. Conversely, explic…

cs.CV2023

AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control

Ruixiang Jiang, Can Wang, Jingbo Zhang +4

Neural implicit fields are powerful for representing 3D scenes and generating high-quality novel views, but it remains challenging to use such implicit representations for creating…

cs.CV2018

Gated Context Aggregation Network for Image Dehazing and Deraining

Dongdong Chen, Mingming He, Qingnan Fan +5

Image dehazing aims to recover the uncorrupted content from a hazy image. Instead of leveraging traditional low-level or handcrafted image priors as the restoration constraints, e.…

cs.CV2026

Lighting in Motion: Spatiotemporal HDR Lighting Estimation

Christophe Bolduc, Julien Philip, Li Ma +3

We present Lighting in Motion (LiMo), a diffusion-based approach to spatiotemporal lighting estimation. LiMo targets both realistic high-frequency detail prediction and accurate il…