papers

Publications (25)

cs.CV2025

The Aging Multiverse: Generating Condition-Aware Facial Aging Tree via Training-Free Diffusion

Bang Gong, Luchao Qi, Jiaye Wu +5

We introduce the Aging Multiverse, a framework for generating multiple plausible facial aging trajectories from a single image, each conditioned on external factors such as environ…

cs.CV2024

Leveraging Near-Field Lighting for Monocular Depth Estimation from Endoscopy Videos

Akshay Paruchuri, Samuel Ehrenstein, Shuxian Wang +4

Monocular depth estimation in endoscopy videos can enable assistive and robotic surgery to obtain better coverage of the organ and detection of various health issues. Despite promi…

cs.CV2025

GLOW: Global Illumination-Aware Inverse Rendering of Indoor Scenes Captured with Dynamic Co-Located Light & Camera

Jiaye Wu, Saeed Hadadan, Geng Lin +4

Inverse rendering of indoor scenes remains challenging due to the ambiguity between reflectance and lighting, exacerbated by inter-reflections among multiple objects. While natural…

cs.CV2024

My3DGen: A Scalable Personalized 3D Generative Model

Luchao Qi, Jiaye Wu, Annie N. Wang +2

In recent years, generative 3D face models (e.g., EG3D) have been developed to tackle the problem of synthesizing photo-realistic faces. However, these models are often unable to c…

cs.CV2026

MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos

Daniel Rho, Jun Myeong Choi, Matthew Thornton +2

Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings,…

cs.CV2026

Understanding Model Behavior in Monocular Polyp Sizing

Xinqi Xiong, Andrea Dunn Beltran, Junmyeong Choi +3

Accurate polyp size stratification guides surveillance decisions, with lesions larger than 5 mm typically requiring closer follow-up. However, monocular colonoscopy lacks a reliabl…

cs.CV2026

Stop Holding Your Breath: CT-Informed Gaussian Splatting for Dynamic Bronchoscopy

Andrea Dunn Beltran, Daniel Rho, Aarav Mehta +5

Bronchoscopic navigation relies on registering endoscopic video to a preoperative CT scan, but respiratory motion deforms the airway by 5-20 mm, creating CT-to-body divergence that…

cs.CV2026

NFL-BA: Near-Field Light Bundle Adjustment for SLAM in Dynamic Lighting

Andrea Dunn Beltran, Daniel Rho, Marc Niethammer +1

Simultaneous Localization and Mapping (SLAM) systems typically assume static, distant illumination; however, many real-world scenarios, such as endoscopy, subterranean robotics, an…

cs.CV2026

HarmoVid: Relightful Video Portrait Harmonization

Jun Myeong Choi, Jae Shin Yoon, Luchao Qi +2

We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illumination intensity (relightful har…

cs.CV2025

Over++: Generative Video Compositing for Layer Interaction Effects

Luchao Qi, Jiaye Wu, Jun Myeong Choi +3

In professional video compositing workflows, artists must manually create environmental interactions-such as shadows, reflections, dust, and splashes-between foreground subjects an…

cs.CV2026

GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures

Patrick Noras, Jun Myeong Choi, Didier Stricker +2

GAINS is a two‑stage inverse‑rendering system that combines Gaussian splatting with foundation‑model priors (depth, normals, segmentation, diffusion) to recover accurate geometry a…

#inverse rendering#gaussian splatting#sparse multi-view#foundation model priors
cs.CV2025

PPS-Ctrl: Controllable Sim-to-Real Translation for Colonoscopy Depth Estimation

Xinqi Xiong, Andrea Dunn Beltran, Jun Myeong Choi +2

Accurate depth estimation enhances endoscopy navigation and diagnostics, but obtaining ground-truth depth in clinical settings is challenging. Synthetic datasets are often used for…

cs.CV2023

Joint Depth Prediction and Semantic Segmentation with Multi-View SAM

Mykhailo Shvets, Dongxu Zhao, Marc Niethammer +2

Multi-task approaches to joint depth and segmentation prediction are well-studied for monocular images. Yet, predictions from a single-view are inherently limited, while multiple v…

cs.CV2024

Personalized Video Relighting With an At-Home Light Stage

Jun Myeong Choi, Max Christman, Roni Sengupta

In this paper, we develop a personalized video relighting algorithm that produces high-quality and temporally consistent relit videos under any pose, expression, and lighting condi…

cs.CV2025

Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering

Noah Frahm, Prakrut Patel, Yue Zhang +3

Large vision-language models (VLMs) have improved embodied question answering (EQA) agents by providing strong semantic priors for open-vocabulary reasoning. However, when used dir…

cs.CV2024

Continual Learning of Personalized Generative Face Models with Experience Replay

Annie N. Wang, Luchao Qi, Roni Sengupta

We introduce a novel continual learning problem: how to sequentially update the weights of a personalized 2D and 3D generative face model as new batches of photos in different appe…

cs.CV2024

ScribbleLight: Single Image Indoor Relighting with Scribbles

Jun Myeong Choi, Annie Wang, Pieter Peers +2

Image-based relighting of indoor rooms creates an immersive virtual understanding of the space, which is useful for interior design, virtual staging, and real estate. Relighting in…

cs.CV2026

ProJo4D: Progressive Joint Optimization for Sparse-View Inverse Physics Estimation

Daniel Rho, Jun Myeong Choi, Biswadip Dey +1

Neural rendering has advanced significantly in 3D reconstruction and novel view synthesis, and integrating physics into these frameworks opens new applications such as physically a…

cs.CV2026

TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake Detection

Xinqi Xiong, Prakrut Patel, Qingyuan Fan +6

The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks…

cs.CV2025

VIN-NBV: A View Introspection Network for Next-Best-View Selection

Noah Frahm, Dongxu Zhao, Andrea Dunn Beltran +4

Next Best View (NBV) algorithms aim to maximize 3D scene acquisition quality using minimal resources, e.g. number of acquisitions, time taken, or distance traversed. Prior methods…

cs.CV2025

MyTimeMachine: Personalized Facial Age Transformation

Luchao Qi, Jiaye Wu, Bang Gong +3

Facial aging is a complex process, highly dependent on multiple factors like gender, ethnicity, lifestyle, etc., making it extremely challenging to learn a global aging prior to pr…

cs.CV2025

EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection

Huaizhi Qu, Ruichen Zhang, Shuqing Luo +5

Recent advances in foundation models have driven remarkable progress in image editing, yet their extension to 3D editing remains underexplored. A natural approach is to replace the…

cs.CV2024

: Neural Deformation Fields for Approximately Diffeomorphic Medical Image Registration

Lin Tian, Hastings Greer, Raúl San José Estépar +2

This work proposes NePhi, a generalizable neural deformation model which results in approximately diffeomorphic transformations. In contrast to the predominant voxel-based transfor…

cs.CV2026

GaNI: Global and Near Field Illumination Aware Neural Inverse Rendering

Jiaye Wu, Saeed Hadadan, Geng Lin +3

In this paper, we present GaNI, a Global and Near-field Illumination-aware neural inverse rendering technique that can reconstruct geometry, albedo, and roughness parameters from i…

cs.CV2024

Structure-preserving Image Translation for Depth Estimation in Colonoscopy Video

Shuxian Wang, Akshay Paruchuri, Zhaoxi Zhang +2

Monocular depth estimation in colonoscopy video aims to overcome the unusual lighting properties of the colonoscopic environment. One of the major challenges in this area is the do…