3 papers
cs.LG2025
Adapting Vision-Language Models for Evaluating World Models
Mariya Hendriksen, Tabish Rashid, David Bignell +5
World models - generative models that simulate environment dynamics conditioned on past observations and actions - are gaining prominence in planning, simulation, and embodied AI.…
cs.CV2025
Fast Autoregressive Video Generation with Diagonal Decoding
Yang Ye, Junliang Guo, Haoyu Wu +5
Autoregressive Transformer models have demonstrated impressive performance in video generation, but their sequential token-by-token decoding process poses a major bottleneck, parti…
cs.LG2024
Scaling Laws for Pre-training Agents and World Models
Tim Pearce, Tabish Rashid, Dave Bignell +3
The performance of embodied agents has been shown to improve by increasing model parameters, dataset size, and compute. This has been demonstrated in domains from robotics to video…