13 papers
TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems
Shunwen Bai, Ziping Ma, Chaoyang Zhang +4
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and all…
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
Dehao Hao, Kaiyi Zhang, Tanghui Jia +10
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. E…
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
Mingxian Lin, Shengju Qian, Yuqi Liu +9
Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per…
SAP: Segment Any 4K Panorama
Lutao Jiang, Zidong Cao, Weikai Chen +14
Promptable instance segmentation is widely adopted in embodied and AR systems, yet the performance of foundation models trained on perspective imagery often degrades on 360° panor…
CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose
Li Jin, Yuchen Yang, Weikai Chen +11
3D learning systems implicitly assume that objects occupy a coherent reference frame. Nonetheless, in practice, every asset arrives with an arbitrary global rotation, and models ar…
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
Li Jin, Weikai Chen, Yujie Wang +7
Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional r…