1 paper
Zixian Guo, Ming Liu, Qilong Wang +4
Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end tra…