activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Peiwen Sun, Shiqiang Lang, Dongming Wu +8

With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications su…

cs.CV2026

DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

Yani Zhang, Dongming Wu, Hao Shi +3

A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pair a detector with a separate gr…

cs.CV2025

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

Dongming Wu, Yanping Fu, Saike Huang +8

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from th…

cs.CV2024

Merlin:Empowering Multimodal LLMs with Foresight Minds

En Yu, Liang Zhao, Yana Wei +8

Humans possess the remarkable ability to foresee the future to a certain extent based on present observations, a skill we term as foresight minds. However, this capability remains…

cs.CV2024

Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?

Yifan Bai, Dongming Wu, Yingfei Liu +8

Rapid advancements in Autonomous Driving (AD) tasks turned a significant shift toward end-to-end fashion, particularly in the utilization of vision-language models (VLMs) that inte…