3 papers
cs.CV2026
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Weiliang Chen, Haowen Sun, Jun Gao +40
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…
cs.CV2026
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
Shangchen Miao, Ningya Feng, Jialong Wu +4
Recent vision-language-action (VLA) models built upon pretrained vision-language models (VLMs) have achieved significant improvements in robotic manipulation. However, current VLAs…
cs.CV2025
Vid2World: Crafting Video Diffusion Models to Interactive World Models
Siqiao Huang, Jialong Wu, Qixing Zhou +2
World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. How…