1 paper
Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche +7
Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, whic…