Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Hierarchical Denoising For Multi-Step Visual Reasoning
Zezhong Qian, Xiaowei Chi, Chak-Wing Mak +9
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in…
cs.CV2026
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16
Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…