2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 2 cited
Mitigating Hallucination in Visual Language Models with Visual Supervision
Zhiyang Chen, Yousong Zhu, Yufei Zhan +4
Large vision-language models (LVLMs) suffer from hallucination a lot, generating responses that apparently contradict to the image content occasionally. The key problem lies in its…
cs.LG2023★ 1 cited
Continual Instruction Tuning for Large Multimodal Models
Jinghan He, Haiyun Guo, Ming Tang +1
Instruction tuning is now a widely adopted approach to aligning large multimodal models (LMMs) to follow human intent. It unifies the data format of vision-language tasks, enabling…