activity
20162023
most citedBURST: A Benchmark for Unifying Object Recognition, Segmentation and Tracking in Video

2 citations · 2 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV202326 cited

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Xi Chen, Xiao Wang, Lucas Beyer +16

This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this s…

cs.CV20231 cited

StoryBench: A Multifaceted Benchmark for Continuous Story Visualization

Emanuele Bugliarello, Hernan Moraldo, Ruben Villegas +7

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst be…

cs.CV2023

Connecting Vision and Language with Video Localized Narratives

Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset +2

We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move th…

cs.CV20222 cited

BURST: A Benchmark for Unifying Object Recognition, Segmentation and Tracking in Video

Ali Athar, Jonathon Luiten, Paul Voigtlaender +4

Multiple existing benchmarks involve tracking and segmenting objects in video e.g., Video Object Segmentation (VOS) and Multi-Object Tracking and Segmentation (MOTS), but there is…

cs.LG2016

RETURNN: The RWTH Extensible Training framework for Universal Recurrent Neural Networks

Patrick Doetsch, Albert Zeyer, Paul Voigtlaender +3

In this work we release our extensible and easily configurable neural network training software. It provides a rich set of functional layers with a particular focus on efficient tr…