9 papers · 1 filter
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Junxiang Xu, Ruisi Wang, Fanyi Pu +49
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be r…
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Yukang Cao, Haozhe Xie, Beichen Wen +13
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touc…
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
Yukang Cao, Haozhe Xie, Fangzhou Hong +4
We present HSImul3R, a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casual captures, including sparse-view images and monocular v…
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
Zixian Liu, Zhaoxi Chen, Liang Pan +1
In recent years, researchers have increasingly been interested in how to enable Multimodal Large Language Models (MLLM) to possess spatial understanding and reasoning capabilities.…
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
Ziang Cao, Fangzhou Hong, Zhaoxi Chen +2
3D modeling is shifting from static visual representations toward physical, articulated assets that can be directly used in simulation and interaction. However, most existing 3D ge…
Scaling Spatial Intelligence with Multimodal Foundation Models
Zhongang Cai, Ruisi Wang, Chenyang Gu +26
Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation m…