activity
20192021
most citedMMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering

10 citations · 10 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV2021

Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using Capsules

Aisha Urooj Khan, Hilde Kuehne, Kevin Duarte +3

The problem of grounding VQA tasks has seen an increased attention in the research community recently, with most attempts usually focusing on solving this task by using pretrained…

cs.CV202010 cited

MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering

Aisha Urooj Khan, Amir Mazaheri, Niels da Vitoria Lobo +1

We present MMFT-BERT(MultiModal Fusion Transformer with BERT encodings), to solve Visual Question Answering (VQA) ensuring individual and combined processing of multiple input moda…

cs.CV2020

Deep Photo Cropper and Enhancer

Aaron Ott, Amir Mazaheri, Niels D. Lobo +1

This paper introduces a new type of image enhancement problem. Compared to traditional image enhancement methods, which mostly deal with pixel-wise modifications of a given photo,…

cs.CV2020

Text Synopsis Generation for Egocentric Videos

Aidean Sharghi, Niels da Vitoria Lobo, Mubarak Shah

Mass utilization of body-worn cameras has led to a huge corpus of available egocentric video. Existing video summarization algorithms can accelerate browsing such videos by selecti…

cs.CV2019

Unsupervised Visual Representation Learning with Increasing Object Shape Bias

Zhibo Wang, Shen Yan, Xiaoyu Zhang +1

(Very early draft)Traditional supervised learning keeps pushing convolution neural network(CNN) achieving state-of-art performance. However, lack of large-scale annotation data is…