1 paper · 1 filter
Bohan Zhou, Zhongbin Zhang, Jiangxing Wang +1
The in-context learning ability of Transformer models has brought new possibilities to visual navigation. In this paper, we focus on the video navigation setting, where an in-conte…