2 citations · 2 across the 3 of their papers we have counts for
5 papers
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Xi Chen, Xiao Wang, Lucas Beyer +16
This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this s…
StoryBench: A Multifaceted Benchmark for Continuous Story Visualization
Emanuele Bugliarello, Hernan Moraldo, Ruben Villegas +7
Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst be…
Connecting Vision and Language with Video Localized Narratives
Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset +2
We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move th…
BURST: A Benchmark for Unifying Object Recognition, Segmentation and Tracking in Video
Ali Athar, Jonathon Luiten, Paul Voigtlaender +4
Multiple existing benchmarks involve tracking and segmenting objects in video e.g., Video Object Segmentation (VOS) and Multi-Object Tracking and Segmentation (MOTS), but there is…
RETURNN: The RWTH Extensible Training framework for Universal Recurrent Neural Networks
Patrick Doetsch, Albert Zeyer, Paul Voigtlaender +3
In this work we release our extensible and easily configurable neural network training software. It provides a rich set of functional layers with a particular focus on efficient tr…