10 citations · 11 across the 2 of their papers we have counts for
4 papers
Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using Capsules
Aisha Urooj Khan, Hilde Kuehne, Kevin Duarte +3
The problem of grounding VQA tasks has seen an increased attention in the research community recently, with most attempts usually focusing on solving this task by using pretrained…
MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering
Aisha Urooj Khan, Amir Mazaheri, Niels da Vitoria Lobo +1
We present MMFT-BERT(MultiModal Fusion Transformer with BERT encodings), to solve Visual Question Answering (VQA) ensuring individual and combined processing of multiple input moda…
Analysis of Hand Segmentation in the Wild
Aisha Urooj Khan, Ali Borji
A large number of works in egocentric vision have concentrated on action and object recognition. Detection and segmentation of hands in first-person videos, however, has less been…
Segmenting Sky Pixels in Images
Cecilia La Place, Aisha Urooj Khan, Ali Borji
Outdoor scene parsing models are often trained on ideal datasets and produce quality results. However, this leads to a discrepancy when applied to the real world. The quality of sc…