From the 1 of 7 linked papers with an AI index.
7 papers
Controlling Embedding Spaces with Text-Conditioned Transformations
Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani +2
Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compre…
Autoregressive Modeling of Film with Applications in Video Montage
Marcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar +5
FilmGPT is an autoregressive transformer trained on a large movie corpus to learn the statistical patterns of film editing and select existing raw shots to create coherent video mo…
Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition
Prajwal Gatti, Simon Jenni, Fabian Caba Heilbron +1
We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on d…
ResidualViT for Efficient Temporally Dense Video Encoding
Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem +2
Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" r…
Discovering Divergent Representations between Text-to-Image Models
Lisa Dunlap, Joseph E. Gonzalez, Trevor Darrell +3
In this paper, we investigate when and how visual representations learned by two different generative models diverge. Given two text-to-image models, our goal is to discover visual…
Improving Personalized Search with Regularized Low-Rank Parameter Updates
Fiona Ryan, Josef Sivic, Fabian Caba Heilbron +3
Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning…