Publications (25)
The Aging Multiverse: Generating Condition-Aware Facial Aging Tree via Training-Free Diffusion
Bang Gong, Luchao Qi, Jiaye Wu +5
We introduce the Aging Multiverse, a framework for generating multiple plausible facial aging trajectories from a single image, each conditioned on external factors such as environ…
Leveraging Near-Field Lighting for Monocular Depth Estimation from Endoscopy Videos
Akshay Paruchuri, Samuel Ehrenstein, Shuxian Wang +4
Monocular depth estimation in endoscopy videos can enable assistive and robotic surgery to obtain better coverage of the organ and detection of various health issues. Despite promi…
GLOW: Global Illumination-Aware Inverse Rendering of Indoor Scenes Captured with Dynamic Co-Located Light & Camera
Jiaye Wu, Saeed Hadadan, Geng Lin +4
Inverse rendering of indoor scenes remains challenging due to the ambiguity between reflectance and lighting, exacerbated by inter-reflections among multiple objects. While natural…
My3DGen: A Scalable Personalized 3D Generative Model
Luchao Qi, Jiaye Wu, Annie N. Wang +2
In recent years, generative 3D face models (e.g., EG3D) have been developed to tackle the problem of synthesizing photo-realistic faces. However, these models are often unable to c…
MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos
Daniel Rho, Jun Myeong Choi, Matthew Thornton +2
Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings,…
Understanding Model Behavior in Monocular Polyp Sizing
Xinqi Xiong, Andrea Dunn Beltran, Junmyeong Choi +3
Accurate polyp size stratification guides surveillance decisions, with lesions larger than 5 mm typically requiring closer follow-up. However, monocular colonoscopy lacks a reliabl…
Stop Holding Your Breath: CT-Informed Gaussian Splatting for Dynamic Bronchoscopy
Andrea Dunn Beltran, Daniel Rho, Aarav Mehta +5
Bronchoscopic navigation relies on registering endoscopic video to a preoperative CT scan, but respiratory motion deforms the airway by 5-20 mm, creating CT-to-body divergence that…
NFL-BA: Near-Field Light Bundle Adjustment for SLAM in Dynamic Lighting
Andrea Dunn Beltran, Daniel Rho, Marc Niethammer +1
Simultaneous Localization and Mapping (SLAM) systems typically assume static, distant illumination; however, many real-world scenarios, such as endoscopy, subterranean robotics, an…
HarmoVid: Relightful Video Portrait Harmonization
Jun Myeong Choi, Jae Shin Yoon, Luchao Qi +2
We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illumination intensity (relightful har…
Over++: Generative Video Compositing for Layer Interaction Effects
Luchao Qi, Jiaye Wu, Jun Myeong Choi +3
In professional video compositing workflows, artists must manually create environmental interactions-such as shadows, reflections, dust, and splashes-between foreground subjects an…
GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures
Patrick Noras, Jun Myeong Choi, Didier Stricker +2
GAINS is a two‑stage inverse‑rendering system that combines Gaussian splatting with foundation‑model priors (depth, normals, segmentation, diffusion) to recover accurate geometry a…
PPS-Ctrl: Controllable Sim-to-Real Translation for Colonoscopy Depth Estimation
Xinqi Xiong, Andrea Dunn Beltran, Jun Myeong Choi +2
Accurate depth estimation enhances endoscopy navigation and diagnostics, but obtaining ground-truth depth in clinical settings is challenging. Synthetic datasets are often used for…
Joint Depth Prediction and Semantic Segmentation with Multi-View SAM
Mykhailo Shvets, Dongxu Zhao, Marc Niethammer +2
Multi-task approaches to joint depth and segmentation prediction are well-studied for monocular images. Yet, predictions from a single-view are inherently limited, while multiple v…
Personalized Video Relighting With an At-Home Light Stage
Jun Myeong Choi, Max Christman, Roni Sengupta
In this paper, we develop a personalized video relighting algorithm that produces high-quality and temporally consistent relit videos under any pose, expression, and lighting condi…
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
Noah Frahm, Prakrut Patel, Yue Zhang +3
Large vision-language models (VLMs) have improved embodied question answering (EQA) agents by providing strong semantic priors for open-vocabulary reasoning. However, when used dir…
Continual Learning of Personalized Generative Face Models with Experience Replay
Annie N. Wang, Luchao Qi, Roni Sengupta
We introduce a novel continual learning problem: how to sequentially update the weights of a personalized 2D and 3D generative face model as new batches of photos in different appe…
ScribbleLight: Single Image Indoor Relighting with Scribbles
Jun Myeong Choi, Annie Wang, Pieter Peers +2
Image-based relighting of indoor rooms creates an immersive virtual understanding of the space, which is useful for interior design, virtual staging, and real estate. Relighting in…
ProJo4D: Progressive Joint Optimization for Sparse-View Inverse Physics Estimation
Daniel Rho, Jun Myeong Choi, Biswadip Dey +1
Neural rendering has advanced significantly in 3D reconstruction and novel view synthesis, and integrating physics into these frameworks opens new applications such as physically a…
TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake Detection
Xinqi Xiong, Prakrut Patel, Qingyuan Fan +6
The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks…
VIN-NBV: A View Introspection Network for Next-Best-View Selection
Noah Frahm, Dongxu Zhao, Andrea Dunn Beltran +4
Next Best View (NBV) algorithms aim to maximize 3D scene acquisition quality using minimal resources, e.g. number of acquisitions, time taken, or distance traversed. Prior methods…
MyTimeMachine: Personalized Facial Age Transformation
Luchao Qi, Jiaye Wu, Bang Gong +3
Facial aging is a complex process, highly dependent on multiple factors like gender, ethnicity, lifestyle, etc., making it extremely challenging to learn a global aging prior to pr…
EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
Huaizhi Qu, Ruichen Zhang, Shuqing Luo +5
Recent advances in foundation models have driven remarkable progress in image editing, yet their extension to 3D editing remains underexplored. A natural approach is to replace the…
: Neural Deformation Fields for Approximately Diffeomorphic Medical Image Registration
Lin Tian, Hastings Greer, Raúl San José Estépar +2
This work proposes NePhi, a generalizable neural deformation model which results in approximately diffeomorphic transformations. In contrast to the predominant voxel-based transfor…
GaNI: Global and Near Field Illumination Aware Neural Inverse Rendering
Jiaye Wu, Saeed Hadadan, Geng Lin +3
In this paper, we present GaNI, a Global and Near-field Illumination-aware neural inverse rendering technique that can reconstruct geometry, albedo, and roughness parameters from i…
Structure-preserving Image Translation for Depth Estimation in Colonoscopy Video
Shuxian Wang, Akshay Paruchuri, Zhaoxi Zhang +2
Monocular depth estimation in colonoscopy video aims to overcome the unusual lighting properties of the colonoscopic environment. One of the major challenges in this area is the do…