1 paper · 1 filter
Neale Ratzlaff, Man Luo, Xin Su +2
Multimodal models typically combine a powerful large language model (LLM) with a vision encoder and are then trained on multimodal data via instruction tuning. While this process a…