1 paper · 1 filter
Shraman Pramanick, Guangxing Han, Rui Hou +6
The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems, unifying various vision-language (VL) tasks by instruction tu…