From the 1 of 35 linked papers with an AI index.
2 citations · 2 across the 14 of their papers we have counts for
15 papers · 1 filter
Hierarchical Denoising For Multi-Step Visual Reasoning
Zezhong Qian, Xiaowei Chi, Chak-Wing Mak +9
The paper introduces HDR, a hierarchical denoising framework for causal video generation that enables multi-step visual reasoning with low-latency streaming, achieving higher succe…
WAM4D: Fast 4D World Action Model via Spatial Register Tokens
Ying Li, Xiaobao Wei, Jiajun Cao +10
World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video o…
PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing
Peng Li, Wangguandong Zheng, Yuan Liu +9
Detailed and photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, full-body reconstruction from a monocular RGB image r…
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16
Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
Zezhong Qian, Xiaowei Chi, Yuming Li +5
Wrist-view observations are crucial for VLA models as they capture fine-grained hand-object interactions that directly enhance manipulation performance. Yet large-scale datasets ra…
Can World Models Benefit VLMs for World Dynamics?
Kevin Zhang, Kuangzhi Ge, Xiaowei Chi +5
Trained on internet-scale video data, generative world models are increasingly recognized as powerful world simulators that can generate consistent and plausible dynamics over stru…