2 papers
cs.CV2026
Video Models Reason Early: Exploiting Plan Commitment for Maze Solving
Kaleb Newman, Tyler Zhu, Olga Russakovsky
Video diffusion models exhibit emergent reasoning capabilities like solving mazes and puzzles, yet little is understood about how they reason during generation. We take a first ste…
cs.CV2024
Do Pre-trained Vision-Language Models Encode Object States?
Kaleb Newman, Shijie Wang, Yuan Zang +2
For a vision-language model (VLM) to understand the physical world, such as cause and effect, a first step is to capture the temporal dynamics of the visual world, for example how…