2 citations · 5 across the 4 of their papers we have counts for
1 paper · 1 filter
Yi-Kai Zhang, Shiyin Lu, Yang Li +7
Multimodal large language models (MLLMs), initiated with a trained LLM, first align images with text and then fine-tune on multimodal mixed inputs. However, the MLLM catastrophical…