1 paper
Yiming Fu, Fangjun Li, Xiujin Liu +7
Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before languag…