8 papers
Apple-: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Runmao Yao, Kairui Hu, Yukang Cao +11
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausi…
Demystifying Video Reasoning
Ruisi Wang, Zhongang Cai, Fanyi Pu +11
Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities. Prior work attributes this to a…
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Haiwen Diao, Penghao Wu, Hanming Deng +55
Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fra…
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
Jing Lin, Ruisi Wang, Junzhe Lu +10
Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization cap…
Scaling Spatial Intelligence with Multimodal Foundation Models
Zhongang Cai, Ruisi Wang, Chenyang Gu +26
Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation m…
A Very Big Video Reasoning Suite
Maijunxian Wang, Ruisi Wang, Juyi Lin +53
Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally c…