5 papers
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
Dean L Slack, G Thomas Hudson, Thomas Winterbottom +1
Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study inves…
Everything is a Video: Unifying Modalities through Next-Frame Prediction
G. Thomas Hudson, Dean Slack, Thomas Winterbottom +4
Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual questi…
The Power of Next-Frame Prediction for Learning Physical Laws
Thomas Winterbottom, G. Thomas Hudson, Daniel Kluvanec +6
Next-frame prediction is a useful and powerful method for modelling and understanding the dynamics of video data. Inspired by the empirical success of causal language modelling and…
RAR-b: Reasoning as Retrieval Benchmark
Chenghao Xiao, G Thomas Hudson, Noura Al Moubayed
Semantic textual similartiy (STS) and information retrieval tasks (IR) tasks have been the two major avenues to record the progress of embedding models in the past few years. Under…
Pixel Sentence Representation Learning
Chenghao Xiao, Zhuoxu Huang, Danlu Chen +7
Pretrained language models are long known to be subpar in capturing sentence and document-level semantics. Though heavily investigated, transferring perturbation-based methods from…