Publications (15)
MVLayoutNet:3D layout reconstruction with multi-view panoramas
Zhihua Hu, Bo Duan, Yanfeng Zhang +2
We present MVLayoutNet, an end-to-end network for holistic 3D reconstruction from multi-view panoramas. Our core contribution is to seamlessly combine learned monocular layout esti…
GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis
Xuqin Wang, Tao Wu, Yanfeng Zhang +5
Recent advances in generative modeling have substantially enhanced novel view synthesis, yet maintaining consistency across viewpoints remains challenging. Diffusion-based models r…
Primitive Graph Learning for Unified Vector Mapping
Lei Wang, Min Dai, Jianan He +2
Large-scale vector mapping is important for transportation, city planning, and survey and census. We propose GraphMapper, a unified framework for end-to-end vector map extraction f…
LADB: Latent Aligned Diffusion Bridges for Semi-Supervised Domain Translation
Xuqin Wang, Tao Wu, Yanfeng Zhang +6
Diffusion models excel at generating high-quality outputs but face challenges in data-scarce domains, where exhaustive retraining or costly paired data are often required. To addre…
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
Fernando Ropero, Erkin Turkoz, Daniel Matos +6
Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approac…
EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention
Yulong Shi, Mingwei Sun, Yongshuai Wang +2
Owing to advancements in deep learning technology, Vision Transformers (ViTs) have demonstrated impressive performance in various computer vision tasks. Nonetheless, ViTs still fac…