activity
20152022
most citedReferring Transformer: A One-step Approach to Multi-task Visual Grounding

73 citations · 189 across the 22 of their papers we have counts for

collaborators

45 papers

cs.LG2022

GraphPNAS: Learning Distribution of Good Neural Architectures via Deep Graph Generative Models

Muchen Li, Jeffrey Yunfan Liu, Leonid Sigal +1

Neural architectures can be naturally viewed as computational graphs. Motivated by this perspective, we, in this paper, study neural architecture search (NAS) through the lens of l…

cs.CV20221 cited

VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge

Sahithya Ravi, Aditya Chinchure, Leonid Sigal +2

There has been a growing interest in solving Visual Question Answering (VQA) tasks that require the model to reason beyond the content present in the image. In this work, we focus…

cs.CV20225 cited

Real-Time Monitoring of User Stress, Heart Rate and Heart Rate Variability on Mobile Devices

Peyman Bateni, Leonid Sigal

Stress is considered to be the epidemic of the 21st-century. Yet, mobile apps cannot directly evaluate the impact of their content and services on user stress. We introduce the Bea…

cs.CV2021

TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation

Tanzila Rahman, Mengyu Yang, Leonid Sigal

The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-moda…

cs.CV202173 cited

Referring Transformer: A One-step Approach to Multi-task Visual Grounding

Muchen Li, Leonid Sigal

As an important step towards visual reasoning, visual grounding (e.g., phrase localization, referring expression comprehension/segmentation) has been widely explored Previous appro…

cs.CV2021

Segmentation-grounded Scene Graph Generation

Siddhesh Khandelwal, Mohammed Suhail, Leonid Sigal

Scene graph generation has emerged as an important problem in computer vision. While scene graphs provide a grounded representation of objects, their locations and relations in an…