163 citations · 272 across the 9 of their papers we have counts for
1 paper · 1 filter
Zhewei Yao, Xiaoxia Wu, Conglong Li +6
Most of the existing multi-modal models, hindered by their incapacity to adeptly manage interleaved image-and-text inputs in multi-image, multi-round dialogues, face substantial co…