1 paper · 1 filter
Chashi Mahiul Islam, Shaeke Salman, Montasir Shams +2
Building on the unprecedented capabilities of large language models for command understanding and zero-shot recognition of multi-modal vision-language transformers, visual language…