1 citations · 1 across the 12 of their papers we have counts for
1 paper · 1 filter
Benno Krojer, Shravan Nayak, Oscar Mañas +4
Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM.…