3 citations · 3 across the 2 of their papers we have counts for
3 papers · 1 filter
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
Dean L Slack, G Thomas Hudson, Thomas Winterbottom +1
Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study inves…
Everything is a Video: Unifying Modalities through Next-Frame Prediction
G. Thomas Hudson, Dean Slack, Thomas Winterbottom +4
Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual questi…
The Power of Next-Frame Prediction for Learning Physical Laws
Thomas Winterbottom, G. Thomas Hudson, Daniel Kluvanec +6
Next-frame prediction is a useful and powerful method for modelling and understanding the dynamics of video data. Inspired by the empirical success of causal language modelling and…