activity
20182021
most citedAttention Guided Semantic Relationship Parsing for Visual Question Answering

1 citations · 1 across the 1 of their papers we have counts for

collaborators

6 papers

cs.CV2021

Recursive Training for Zero-Shot Semantic Segmentation

Ce Wang, Moshiur Farazi, Nick Barnes

General purpose semantic segmentation relies on a backbone CNN network to extract discriminative features that help classify each image pixel into a 'seen' object class (ie., the o…

cs.CV2021

Efficient Two-Stream Network for Violence Detection Using Separable Convolutional LSTM

Zahidul Islam, Mohammad Rukonuzzaman, Raiyan Ahmed +2

Automatically detecting violence from surveillance footage is a subset of activity recognition that deserves special attention because of its wide applicability in unmanned securit…

cs.CV2020

Rethinking conditional GAN training: An approach using geometrically structured latent manifolds

Sameera Ramasinghe, Moshiur Farazi, Salman Khan +2

Conditional GANs (cGAN), in their rudimentary form, suffer from critical drawbacks such as the lack of diversity in generated outputs and distortion between the latent and output m…

cs.CV20201 cited

Attention Guided Semantic Relationship Parsing for Visual Question Answering

Moshiur Farazi, Salman Khan, Nick Barnes

Humans explain inter-object relationships with semantic labels that demonstrate a high-level understanding required to perform complex Vision-Language tasks such as Visual Question…

cs.CV2020

Accuracy vs. Complexity: A Trade-off in Visual Question Answering Models

Moshiur R. Farazi, Salman H. Khan, Nick Barnes

Visual Question Answering (VQA) has emerged as a Visual Turing Test to validate the reasoning ability of AI agents. The pivot to existing VQA models is the joint embedding that is…

cs.CV2018

From Known to the Unknown: Transferring Knowledge to Answer Questions about Novel Visual and Semantic Concepts

Moshiur R Farazi, Salman H Khan, Nick Barnes

Current Visual Question Answering (VQA) systems can answer intelligent questions about `Known' visual content. However, their performance drops significantly when questions about v…