2 papers
cs.CV2023
Cloud-Device Collaborative Learning for Multimodal Large Language Models
Guanqun Wang, Jiaming Liu, Chenxuan Li +8
The burgeoning field of Multimodal Large Language Models (MLLMs) has exhibited remarkable performance in diverse tasks such as captioning, commonsense reasoning, and visual scene u…
cs.CV2023
MChat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
Xiaowei Chi, Junbo Qi, Rongyu Zhang +3
While current LLM chatbots like GPT-4V bridge the gap between human instructions and visual representations to enable text-image generations, they still lack efficient alignment me…