1 paper · 1 filter
Keunwoo Peter Yu, Joyce Chai
Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. However, real-time interactive environments imp…