1 paper · 1 filter
Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu +1
Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmenta…