1 paper · 1 filter
Clement Neo, Luke Ong, Philip Torr +3
Vision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA…