1 paper · 1 filter
Zihao Zheng, Zhihao Mao, Xingyue Zhou +9
Vision-and-Language Navigation (VLN) increasingly relies on large vision-language models, but their inference cost conflicts with real-time deployment. Token caching is a promising…