196 citations · 419 across the 38 of their papers we have counts for
4 papers · 1 filter
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
Dahun Kim, Anelia Angelova
We propose Context-Adaptive Multi-Prompt Embedding, a novel approach to enrich semantic representations in vision-language contrastive learning. Unlike standard CLIP-style models t…
Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs
Michael Gygli, Mohammad Norouzi, Anelia Angelova
We approach structured output prediction by optimizing a deep value network (DVN) to precisely estimate the task loss on different output configurations for a given input. Once the…
Improved generator objectives for GANs
Ben Poole, Alexander A. Alemi, Jascha Sohl-Dickstein +1
We present a framework to understand GAN training as alternating density ratio estimation and approximate divergence minimization. This provides an interpretation for the mismatche…
Geometry-Based Next Frame Prediction from Monocular Video
Reza Mahjourian, Martin Wicke, Anelia Angelova
We consider the problem of next frame prediction from video input. A recurrent convolutional neural network is trained to predict depth from monocular video input, which, along wit…