1 paper
Ruirui Gao, Emily Johnson, Bowen Tan +1
Large Vision-Language Models (LVLMs) hold immense potential for complex multimodal instruction following, yet their development is often hindered by the high cost and inconsistency…