10 citations · 10 across the 2 of their papers we have counts for
5 papers
Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using Capsules
Aisha Urooj Khan, Hilde Kuehne, Kevin Duarte +3
The problem of grounding VQA tasks has seen an increased attention in the research community recently, with most attempts usually focusing on solving this task by using pretrained…
MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering
Aisha Urooj Khan, Amir Mazaheri, Niels da Vitoria Lobo +1
We present MMFT-BERT(MultiModal Fusion Transformer with BERT encodings), to solve Visual Question Answering (VQA) ensuring individual and combined processing of multiple input moda…
Deep Photo Cropper and Enhancer
Aaron Ott, Amir Mazaheri, Niels D. Lobo +1
This paper introduces a new type of image enhancement problem. Compared to traditional image enhancement methods, which mostly deal with pixel-wise modifications of a given photo,…
Text Synopsis Generation for Egocentric Videos
Aidean Sharghi, Niels da Vitoria Lobo, Mubarak Shah
Mass utilization of body-worn cameras has led to a huge corpus of available egocentric video. Existing video summarization algorithms can accelerate browsing such videos by selecti…
Unsupervised Visual Representation Learning with Increasing Object Shape Bias
Zhibo Wang, Shen Yan, Xiaoyu Zhang +1
(Very early draft)Traditional supervised learning keeps pushing convolution neural network(CNN) achieving state-of-art performance. However, lack of large-scale annotation data is…