55 citations · 256 across the 65 of their papers we have counts for
24 papers · 1 filter
AdVerb: Visually Guided Audio Dereverberation
Sanjoy Chowdhury, Sreyan Ghosh, Subhrajyoti Dasgupta +3
We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverber…
Human Trajectory Forecasting with Explainable Behavioral Uncertainty
Jiangbei Yue, Dinesh Manocha, He Wang
Human trajectory forecasting helps to understand and predict human behaviors, enabling applications from social robots to self-driving cars, and therefore has been heavily investig…
STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded Scenes
Peishan Cong, Xinge Zhu, Feng Qiao +7
Accurately detecting and tracking pedestrians in 3D space is challenging due to large variations in rotations, poses and scales. The situation becomes even worse for dense crowds w…
3MASSIV: Multilingual, Multimodal and Multi-Aspect dataset of Social Media Short Videos
Vikram Gupta, Trisha Mittal, Puneet Mathur +5
We present 3MASSIV, a multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj. 3MASSIV…
SelfTune: Metrically Scaled Monocular Depth Estimation through Self-Supervised Learning
Jaehoon Choi, Dongki Jung, Yonghan Lee +3
Monocular depth estimation in the wild inherently predicts depth up to an unknown scale. To resolve scale ambiguity issue, we present a learning algorithm that leverages monocular…
Active Learning of Neural Collision Handler for Complex 3D Mesh Deformations
Qingyang Tan, Zherong Pan, Breannan Smith +2
We present a robust learning algorithm to detect and handle collisions in 3D deforming meshes. Our collision detector is represented as a bilevel deep autoencoder with an attention…