1 citations · 1 across the 4 of their papers we have counts for
5 papers · 1 filter
Segmenting Collision Sound Sources in Egocentric Videos
Kranti Kumar Parida, Omar Emara, Hazel Doughty +1
Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sou…
HD-EPIC: A Highly-Detailed Egocentric Video Dataset
Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha +16
We present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe…
Beyond Image to Depth: Improving Depth Prediction using Echoes
Kranti Kumar Parida, Siddharth Srivastava, Gaurav Sharma
We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echo…
AVGZSLNet: Audio-Visual Generalized Zero-Shot Learning by Reconstructing Label Features from Multi-Modal Embeddings
Pratik Mazumder, Pravendra Singh, Kranti Kumar Parida +1
In this paper, we propose a novel approach for generalized zero-shot learning in a multi-modal setting, where we have novel classes of audio/video during testing that are not seen…
Coordinated Joint Multimodal Embeddings for Generalized Audio-Visual Zeroshot Classification and Retrieval of Videos
Kranti Kumar Parida, Neeraj Matiyali, Tanaya Guha +1
We present an audio-visual multimodal approach for the task of zeroshot learning (ZSL) for classification and retrieval of videos. ZSL has been studied extensively in the recent pa…