2 citations · 2 across the 10 of their papers we have counts for
11 papers · 1 filter
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
Peng Ling, Yingda Yin, Lingting Zhu +5
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational…
Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs
Kai Li, Lutao Jiang, Zhenyang Li +10
Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room…
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
Dehao Hao, Kaiyi Zhang, Tanghui Jia +10
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. E…
SAP: Segment Any 4K Panorama
Lutao Jiang, Zidong Cao, Weikai Chen +14
Promptable instance segmentation is widely adopted in embodied and AR systems, yet the performance of foundation models trained on perspective imagery often degrades on 360° panora…
CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose
Li Jin, Yuchen Yang, Weikai Chen +11
3D learning systems implicitly assume that objects occupy a coherent reference frame. Nonetheless, in practice, every asset arrives with an arbitrary global rotation, and models ar…
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
Li Jin, Weikai Chen, Yujie Wang +7
Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional r…