8 papers
TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting
Zhi Chen, Yuxuan Wang, Jialong Wu +5
High-quality time series forecasting is pivotal for real-world decision-making. However, traditional point-wise metrics often fail to reveal complex temporal patterns and align poo…
Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models
Yu Zhang, Xingzhuo Guo, Haoran Xu +2
Diffusion and flow-based models have enabled significant progress in generation tasks across various modalities and have recently found applications in predictive learning. However…
Vid2World: Crafting Video Diffusion Models to Interactive World Models
Siqiao Huang, Jialong Wu, Qixing Zhou +2
World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. How…
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
Shangchen Miao, Ningya Feng, Jialong Wu +4
Recent vision-language-action (VLA) models built upon pretrained vision-language models (VLMs) have achieved significant improvements in robotic manipulation. However, current VLAs…
Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
Jialong Wu, Xiaoying Zhang, Hongyi Yuan +7
Humans construct internal world models and reason by manipulating the concepts within these models. Recent advances in AI, particularly chain-of-thought (CoT) reasoning, approximat…
RLVR-World: Training World Models with Reinforcement Learning
Jialong Wu, Shaofeng Yin, Ningya Feng +1
World models predict state transitions in response to actions and are increasingly developed across diverse modalities. However, standard training objectives such as maximum likeli…