1 paper · 1 filter
Thomas Heap, Laurence Aitchison, Emma Cahill +1
We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-t…