1 paper
Serwan Jassim, Mario Holubar, Annika Richter +3
This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This…