1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Moritz Böhle, Amélie Royer, Juliette Marrie +2
Vision-language models (VLMs) are commonly trained by directly inserting image tokens from a pretrained vision encoder into the text stream of a language model. This allows text an…