10 papers
TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting
Zhi Chen, Yuxuan Wang, Jialong Wu +5
High-quality time series forecasting is pivotal for real-world decision-making. However, traditional point-wise metrics often fail to reveal complex temporal patterns and align poo…
Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics
Jiahua Wang, Leqi Zheng, Jialong Wu +2
World models simulate environmental dynamics to enable embodied agents to plan and reason about future states. While real-world perception is inherently multimodal, existing approa…
CompilerDream: Learning a Compiler World Model for General Code Optimization
Chaoyi Deng, Jialong Wu, Ningya Feng +2
Effective code optimization in compilers is crucial for computer and software engineering. The success of these optimizations primarily depends on the selection and ordering of the…
Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models
Yu Zhang, Xingzhuo Guo, Haoran Xu +2
Diffusion and flow-based models have enabled significant progress in generation tasks across various modalities and have recently found applications in predictive learning. However…
Vid2World: Crafting Video Diffusion Models to Interactive World Models
Siqiao Huang, Jialong Wu, Qixing Zhou +2
World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. How…
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
Shangchen Miao, Ningya Feng, Jialong Wu +4
Recent vision-language-action (VLA) models built upon pretrained vision-language models (VLMs) have achieved significant improvements in robotic manipulation. However, current VLAs…