2 citations · 3 across the 18 of their papers we have counts for
Showing 2025 · cs.CVShow all
3 papers · 2 filters
cs.CV2025
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
Hokin Deng
We show that video generation models could reason now. Testing on tasks such as chess, maze, Sudoku, mental rotation, and Raven's Matrices, leading models such as Sora-2 achieve si…
cs.CV2025
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
Zory Zhang, Pinyuan Feng, Bingyang Wang +7
Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation st…
cs.CV2025
Probing Perceptual Constancy in Large Vision-Language Models
Haoran Sun, Bingyang Wang, Suyang Yu +14
Perceptual constancy is the ability to maintain stable perceptions of objects despite changes in sensory input, such as variations in distance, angle, or lighting. This ability is…