most citedVLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation

2 citations · 2 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2025

LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models

Pranav Saxena, Avigyan Bhattacharya, Ji Zhang +1

Referential grounding in outdoor driving scenes is challenging due to large scene variability, many visually similar objects, and dynamic elements that complicate resolving natural…

cs.RO2025

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

Haokun Zhu, Zongtai Li, Zhixuan Liu +4

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs t…

cs.CV2025

SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models

Nader Zantout, Haochen Zhang, Pujith Kachana +4

Interpreting object-referential language and grounding objects in 3D with spatial relations and attributes is essential for robots operating alongside humans. However, this task is…

cs.CV2025

IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes

Haochen Zhang, Nader Zantout, Pujith Kachana +2

With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can…

cs.RO20242 cited

VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation

Haochen Zhang, Nader Zantout, Pujith Kachana +3

With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embo…