From the 1 of 6 linked papers with an AI index.
6 papers
STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding
Mingwei Xing, Xinliang Wang, Yifeng Shi
Constructing a unified 3D scene understanding model has long been hindered by the topological discrepancies across sensor modalities. While applying the Mixture-of-Experts (MoE) ar…
Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
Xinliang Wang, Yijie Kang, Zhenyu Wu +1
Sat2RealCity is a framework that generates 3D urban models from satellite images by grounding object-level 3D generative priors to real-world geographic locations and allowing cont…
AdaptSplat: Adapting Vision Foundation Models for Feed-Forward 3D Gaussian Splatting
Mingwei Xing, Xinliang Wang, Yifeng Shi
This work explores a simple yet powerful lightweight adapter design for feed-forward 3D Gaussian Splatting (3DGS). Existing methods typically apply complex, architecture-specific d…
ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models
Xinliang Wang, Yifeng Shi, Zhenyu Wu
3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constraints. Current generative re…
DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts
Mingwei Xing, Xinliang Wang, Yifeng Shi
Constructing a unified 3D scene understanding model has long been hindered by the significant topological discrepancies across different sensor modalities. While applying the Mixtu…
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
Haoran Lou, Chunxiao Fan, Ziyan Liu +2
The architecture of multimodal large language models (MLLMs) commonly connects a vision encoder, often based on CLIP-ViT, to a large language model. While CLIP-ViT works well for c…