5 papers
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
Hokin Deng
We show that video generation models could reason now. Testing on tasks such as chess, maze, Sudoku, mental rotation, and Raven's Matrices, leading models such as Sora-2 achieve si…
Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
Dezhi Luo, Qingying Gao, Hokin Deng
Spatial world models, representations that support flexible reasoning about spatial relations, are central to developing computational models that could operate in the physical wor…
Proceedings of 1st Workshop on Advancing Artificial Intelligence through Theory of Mind
Mouad Abrini, Omri Abend, Dina Acklin +105
This volume includes a selection of papers presented at the Workshop on Advancing Artificial Intelligence through Theory of Mind held at AAAI 2025 in Philadelphia US on 3rd March 2…
The Philosophical Foundations of Growing AI Like A Child
Dezhi Luo, Yijiang Li, Hokin Deng
Despite excelling in high-level reasoning, current language models lack robustness in real-world scenarios and perform poorly on fundamental problem-solving tasks that are intuitiv…
Probing Perceptual Constancy in Large Vision-Language Models
Haoran Sun, Bingyang Wang, Suyang Yu +14
Perceptual constancy is the ability to maintain stable perceptions of objects despite changes in sensory input, such as variations in distance, angle, or lighting. This ability is…