Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
VideoGameBench: Can Vision-Language Models complete popular video games?
Alex L. Zhang, Thomas L. Griffiths, Karthik R. Narasimhan +1
Vision-language models (VLMs) have achieved strong results on coding and math benchmarks that are challenging for humans, yet their ability to perform tasks that come naturally to…
cs.AI2025
Cognitive Foundations for Reasoning and Their Manifestation in LLMs
Priyanka Kargupta, Shuyue Stella Li, Haocheng Wang +9
Large language models (LLMs) solve complex problems yet fail on simpler variants, suggesting they achieve correct outputs through mechanisms fundamentally different from human reas…