4 citations · 7 across the 8 of their papers we have counts for
4 papers · 2 filters
SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-world Object Detector
Shuailei Ma, Yuefeng Wang, Ying Wei +4
In this paper, we attempt to specialize the VLM model for OWOD tasks by distilling its open-world knowledge into a language-agnostic detector. Surprisingly, we observe that the com…
Contrastive Vision-Language Alignment Makes Efficient Instruction Learner
Lizhao Liu, Xinyu Sun, Tianhang Xiang +3
We study the task of extending the large language model (LLM) into a vision-language instruction-following model. This task is crucial but challenging since the LLM is trained on t…
FGPrompt: Fine-grained Goal Prompting for Image-goal Navigation
Xinyu Sun, Peihao Chen, Jugang Fan +3
Learning to navigate to an image-specified goal is an important but challenging task for autonomous systems. The agent is required to reason the goal location from where a picture…
Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models
Peihao Chen, Xinyu Sun, Hongyan Zhi +5
We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by langu…