1 paper
Pulkit Gera, Faegheh Sardari, Asmar Nadeem +4
Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motio…