1 paper
L'ea Dubois, Klaus Schmidt, Chengyu Wang +3
Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation r…