268 citations · 276 across the 4 of their papers we have counts for
1 paper · 1 filter
Yanshu Li, Jiaqian Li, Kuai Yu +4
Large vision-language models (LVLMs) have demonstrated strong general multimodal capability and are increasingly deployed in downstream systems. This trend has driven growing inter…