most citedBridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill

4 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.RO20234 cited

Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill

Wenzhe Cai, Siyuan Huang, Guangran Cheng +4

Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first…

cs.RO20231 cited

Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions

Yuxing Long, Xiaoqi Li, Wenzhe Cai +1

Visual language navigation (VLN) is an embodied task demanding a wide range of skills encompassing understanding, perception, and planning. For such a multifaceted challenge, previ…

cs.CL2023

VDialogUE: A Unified Evaluation Benchmark for Visually-grounded Dialogue

Yunshui Li, Binyuan Hui, Zhaochao Yin +6

Visually-grounded dialog systems, which integrate multiple modes of communication such as text and visual inputs, have become an increasingly popular area of investigation. However…

cs.CV2023

Whether you can locate or not? Interactive Referring Expression Generation

Fulong Ye, Yuxing Long, Fangxiang Feng +1

Referring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension…

cs.IR2023

Multimodal Recommendation Dialog with Subjective Preference: A New Challenge and Benchmark

Yuxing Long, Binyuan Hui, Caixia Yuan1 +3

Existing multimodal task-oriented dialog data fails to demonstrate the diverse expressions of user subjective preferences and recommendation acts in the real-life shopping scenario…