1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yifan Zhong, Fengshuo Bai, Shaofei Cai +11
The remarkable advancements of vision and language foundation models in multimodal understanding, reasoning, and generation has sparked growing efforts to extend such intelligence…