2 citations · 2 across the 2 of their papers we have counts for
Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
FAST-EQA: Efficient Embodied Question Answering with Global and Local Region Relevancy
Haochen Zhang, Nirav Savaliya, Faizan Siddiqui +1
Embodied Question Answering (EQA) combines visual scene understanding, goal-directed exploration, spatial and temporal reasoning under partial observability. A central challenge is…
cs.RO2024★ 2 cited
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
Haochen Zhang, Nader Zantout, Pujith Kachana +3
With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embo…