1 paper
Viorica Pătrăucean, Lucas Smaira, Ankush Gupta +21
We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GP…