1 paper · 1 filter
Ce Zhang, Zifu Wan, Zhehan Kan +7
While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not alig…