1 paper
Tri Cao, Khoi Le, Thong Nguyen +7
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure i…