4 citations · 4 across the 12 of their papers we have counts for
1 paper · 1 filter
Wanpeng Zhang, Zilong Xie, Yicheng Feng +4
Multimodal Large Language Models have made significant strides in integrating visual and textual information, yet they often struggle with effectively aligning these modalities. We…