1 paper · 1 filter
Haochen Zhang, Nader Zantout, Pujith Kachana +3
With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embo…