1 paper
Tobia Poppi, Burak Uzkent, Amanmeet Garg +7
Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitiga…