activity
20172022
most citedHigh-Order Attention Models for Visual Question Answering

45 citations · 46 across the 3 of their papers we have counts for

collaborators

8 papers

cs.CV2022

Describing Sets of Images with Textual-PCA

Oded Hupert, Idan Schwartz, Lior Wolf

We seek to semantically describe a set of images, capturing both the attributes of single images and the variations within the set. Our procedure is analogous to Principle Componen…

cs.LG20211 cited

Perceptual Score: What Data Modalities Does Your Model Perceive?

Itai Gat, Idan Schwartz, Alexander Schwing

Machine learning advances in the last decade have relied significantly on large-scale datasets that continue to grow in size. Increasingly, those datasets also contain different da…

cs.CV2021

Video and Text Matching with Conditioned Embeddings

Ameen Ali, Idan Schwartz, Tamir Hazan +1

We present a method for matching a text sentence from a given corpus to a given video clip and vice versa. Traditionally video and text matching is done by learning a shared embedd…

cs.AI2021

Ensemble of MRR and NDCG models for Visual Dialog

Idan Schwartz

Assessing an AI agent that can converse in human language and understand visual content is challenging. Generation metrics, such as BLEU scores favor correct syntax over semantics.…

cs.CV2020

Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies

Itai Gat, Idan Schwartz, Alexander Schwing +1

Many recent datasets contain a variety of different data modalities, for instance, image, question, and answer data in visual question answering (VQA). When training deep net class…

cs.CV2019

A Simple Baseline for Audio-Visual Scene-Aware Dialog

Idan Schwartz, Alexander Schwing, Tamir Hazan

The recently proposed audio-visual scene-aware dialog task paves the way to a more data-driven way of learning virtual assistants, smart speakers and car navigation systems. Howeve…