3 papers
cs.CV2025
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
Dean L Slack, G Thomas Hudson, Thomas Winterbottom +1
Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study inves…
cs.CL2025
Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
Dean L. Slack, Noura Al Moubayed
Although large language models excel across many tasks, they can memorise training data and thereby expose private or copyrighted text. Most defences target the pre-training stage,…
cs.CV2024
Everything is a Video: Unifying Modalities through Next-Frame Prediction
G. Thomas Hudson, Dean Slack, Thomas Winterbottom +4
Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual questi…