5 papers · 1 filter
SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
Peiwen Sun, Shiqiang Lang, Dongming Wu +8
With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications su…
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
Yani Zhang, Dongming Wu, Hao Shi +3
A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pair a detector with a separate gr…
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
Dongming Wu, Yanping Fu, Saike Huang +8
General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from th…
Merlin:Empowering Multimodal LLMs with Foresight Minds
En Yu, Liang Zhao, Yana Wei +8
Humans possess the remarkable ability to foresee the future to a certain extent based on present observations, a skill we term as foresight minds. However, this capability remains…
Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
Yifan Bai, Dongming Wu, Yingfei Liu +8
Rapid advancements in Autonomous Driving (AD) tasks turned a significant shift toward end-to-end fashion, particularly in the utilization of vision-language models (VLMs) that inte…