1 paper
Zhaoyang Li, Zhan Ling, Yuchen Zhou +3
Large Vision-Language Models (LVLMs) excel at captioning, visual question answering, and robotics by combining vision and language, yet they often miss obvious objects or hallucina…