1 citations · 1 across the 6 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")
Alvin Wei Ming Tan, David Cardinal, Tania Lorido-Botran +3
Given the inherently multimodal nature of human experience, vision-language models (VLMs) hold substantial promise for modeling human cognition as it grows and develops with experi…
cs.LG2026
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
Dongyan Lin, Phillip Rust, Angel Villar Corrales +19
Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research sugges…