works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding

Mingwei Xing, Xinliang Wang, Yifeng Shi

Constructing a unified 3D scene understanding model has long been hindered by the topological discrepancies across sensor modalities. While applying the Mixture-of-Experts (MoE) ar…

cs.CV2026

Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery

Xinliang Wang, Yijie Kang, Zhenyu Wu +1

Sat2RealCity is a framework that generates 3D urban models from satellite images by grounding object-level 3D generative priors to real-world geographic locations and allowing cont…

cs.CV2026

AdaptSplat: Adapting Vision Foundation Models for Feed-Forward 3D Gaussian Splatting

Mingwei Xing, Xinliang Wang, Yifeng Shi

This work explores a simple yet powerful lightweight adapter design for feed-forward 3D Gaussian Splatting (3DGS). Existing methods typically apply complex, architecture-specific d…

cs.CV2026

ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models

Xinliang Wang, Yifeng Shi, Zhenyu Wu

3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constraints. Current generative re…

cs.CV2026

DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts

Mingwei Xing, Xinliang Wang, Yifeng Shi

Constructing a unified 3D scene understanding model has long been hindered by the significant topological discrepancies across different sensor modalities. While applying the Mixtu…

cs.CV2025

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

Haoran Lou, Chunxiao Fan, Ziyan Liu +2

The architecture of multimodal large language models (MLLMs) commonly connects a vision encoder, often based on CLIP-ViT, to a large language model. While CLIP-ViT works well for c…