collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling

Kaiyi Zhang, Zhihao Liang, Haolin Liu +8

Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active surface cells increase. We i…

cs.CV2026

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

Dehao Hao, Kaiyi Zhang, Tanghui Jia +10

High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. E…

cs.CV2026

Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps

Xiangjun Gao, Zhensong Zhang, Dave Zhenyu Chen +4

We propose Map2Thought, a framework that enables explicit and interpretable spatial reasoning for 3D VLMs. The framework is grounded in two key components: Metric Cognitive Map (Me…

cs.CV2025

Complex-Valued 2D Gaussian Representation for Computer-Generated Holography

Yicheng Zhan, Xiangjun Gao, Long Quan +1

Complex-valued Gaussian primitives have recently been explored for representing holographic radiance fields in 3D novel view synthesis. In this work, we extend this line of researc…

cs.CV2024

DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Wenbo Hu, Xiangjun Gao, Xiaoyu Li +5

Estimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an…

cs.CV2023

ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent Synthesis

Xiangjun Gao, Xiaoyu Li, Chaopeng Zhang +4

In this work, we propose a method to address the challenge of rendering a 3D human from a single image in a free-view manner. Some existing approaches could achieve this by using g…