252 citations · 1k across the 36 of their papers we have counts for
4 papers · 1 filter
TRecViT: A Recurrent Video Transformer
Viorica Pătrăucean, Xu Owen He, Joseph Heyward +10
We propose a novel block for \emph{causal} video modelling. It relies on a time-space-channel factorisation with dedicated blocks for each dimension: gated linear recurrent units (…
Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Viorica Pătrăucean, Lucas Smaira, Ankush Gupta +21
We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GP…
Massively Parallel Video Networks
Joao Carreira, Viorica Patraucean, Laurent Mazare +2
We introduce a class of causal video understanding models that aims to improve efficiency of video processing by maximising throughput, minimising latency, and reducing the number…
Recursive Recurrent Nets with Attention Modeling for OCR in the Wild
Chen-Yu Lee, Simon Osindero
We present recursive recurrent neural networks with attention modeling (RAM) for lexicon-free optical character recognition in natural scene images. The primary advantages of t…