1 citations · 1 across the 3 of their papers we have counts for
6 papers
TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation
Tanzila Rahman, Mengyu Yang, Leonid Sigal
The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-moda…
Weakly-supervised Audio-visual Sound Source Detection and Separation
Tanzila Rahman, Leonid Sigal
Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from arti…
An Improved Attention for Visual Question Answering
Tanzila Rahman, Shih-Han Chou, Leonid Sigal +1
We consider the problem of Visual Question Answering (VQA). Given an image and a free-form, open-ended, question, expressed in natural language, the goal of VQA system is to provid…
Watch, Listen and Tell: Multi-modal Weakly Supervised Dense Event Captioning
Tanzila Rahman, Bicheng Xu, Leonid Sigal
Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from lang…
Convolutional Temporal Attention Model for Video-based Person Re-identification
Tanzila Rahman, Mrigank Rochan, Yang Wang
The goal of video-based person re-identification is to match two input videos, so that the distance of the two videos is small if two videos contain the same person. A common appro…
Video-based Person Re-identification Using Spatial-Temporal Attention Networks
Shivansh Rao, Tanzila Rahman, Mrigank Rochan +1
We consider the problem of video-based person re-identification. The goal is to identify a person from videos captured under different cameras. In this paper, we propose an efficie…