29 citations · 50 across the 25 of their papers we have counts for
29 papers · 1 filter
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
Peng Ling, Yingda Yin, Lingting Zhu +5
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational…
Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs
Kai Li, Lutao Jiang, Zhenyang Li +10
Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room…
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
Dehao Hao, Kaiyi Zhang, Tanghui Jia +10
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. E…
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Zijie Wang, Wei Zhang, Weiming Zhang +4
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth f…
PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments
Zhenyang Li, Lutao Jiang, Yizhou Zhao +4
Reconstructing realistic, physically plausible garments from a single image remains a fundamental challenge. Template-free methods capture surface geometry but lack explicit sewing…
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
Zhijia Liang, Jiaming Li, Weikai Chen +3
Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal i…