activity
20162021
most citedObject Detection with a Unified Label Space from Multiple Datasets

3 citations · 4 across the 5 of their papers we have counts for

collaborators

10 papers

cs.SD20211 cited

Depth Infused Binaural Audio Generation using Hierarchical Cross-Modal Attention

Kranti Kumar Parida, Siddharth Srivastava, Neeraj Matiyali +1

Binaural audio gives the listener the feeling of being in the recording place and enhances the immersive experience if coupled with AR/VR. But the problem with binaural audio recor…

cs.CV2021

Exploiting Local Geometry for Feature and Graph Construction for Better 3D Point Cloud Processing with Graph Neural Networks

Siddharth Srivastava, Gaurav Sharma

We propose simple yet effective improvements in point representations and local neighborhood graph construction within the general framework of graph neural networks (GNNs) for 3D…

cs.CV20203 cited

Object Detection with a Unified Label Space from Multiple Datasets

Xiangyun Zhao, Samuel Schulter, Gaurav Sharma +3

Given multiple datasets with different label spaces, the goal of this work is to train a single object detector predicting over the union of all the label spaces. The practical ben…

cs.CV2019

Coordinated Joint Multimodal Embeddings for Generalized Audio-Visual Zeroshot Classification and Retrieval of Videos

Kranti Kumar Parida, Neeraj Matiyali, Tanaya Guha +1

We present an audio-visual multimodal approach for the task of zeroshot learning (ZSL) for classification and retrieval of videos. ZSL has been studied extensively in the recent pa…

cs.CV2019

Video Person Re-Identification using Learned Clip Similarity Aggregation

Neeraj Matiyali, Gaurav Sharma

We address the challenging task of video-based person re-identification. Recent works have shown that splitting the video sequences into clips and then aggregating clip based simil…

cs.CV2019

Learning 2D to 3D Lifting for Object Detection in 3D for Autonomous Vehicles

Siddharth Srivastava, Frederic Jurie, Gaurav Sharma

We address the problem of 3D object detection from 2D monocular images in autonomous driving scenarios. We propose to lift the 2D images to 3D representations using learned neural…