collaborators

7 papers

cs.CV2025

Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction

Youming Deng, Wenqi Xian, Guandao Yang +4

In this paper, we present a self-calibrating framework that jointly optimizes camera parameters, lens distortion and 3D Gaussian representations, enabling accurate and efficient sc…

cs.CV2025

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images

Boyang Deng, Songyou Peng, Kyle Genova +4

We present a system using Multimodal LLMs (MLLMs) to analyze a large database with tens of millions of images captured at different times, with the aim of discovering patterns in t…

cs.LG2025

Gaussian Mixture Flow Matching Models

Hansheng Chen, Kai Zhang, Hao Tan +5

Diffusion models approximate the denoising distribution as a Gaussian and predict its mean, whereas flow matching models reparameterize the Gaussian mean as flow velocity. However,…

cs.GR2025

GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modeling

Yang Zheng, Menglei Chai, Delio Vicini +5

We present GroomLight, a novel method for relightable hair appearance modeling from multi-view images. Existing hair capture methods struggle to balance photorealistic rendering wi…

cs.CV2025

FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement

Ian Huang, Yanan Bao, Karen Truong +4

Scene generation with 3D assets presents a complex challenge, requiring both high-level semantic understanding and low-level geometric reasoning. While Multimodal Large Language Mo…

cs.CV2025

Synthesizing 3D Abstractions by Inverting Procedural Buildings with Transformers

Maximilian Dax, Jordi Berbel, Jan Stria +2

We generate abstractions of buildings, reflecting the essential aspects of their geometry and structure, by learning to invert procedural models. We first build a dataset of abstra…