88 citations · 204 across the 28 of their papers we have counts for
8 papers · 1 filter
Panoramic Video Salient Object Detection with Ambisonic Audio Guidance
Xiang Li, Haoyuan Cao, Shijie Zhao +3
Video salient object detection (VSOD), as a fundamental computer vision problem, has been extensively discussed in the last decade. However, all existing works focus on addressing…
AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
Khoa Vo, Sang Truong, Kashu Yamazaki +3
Temporal action proposal generation (TAPG) is a challenging task, which requires localizing action intervals in an untrimmed video. Intuitively, we as humans, perceive an action th…
Point3D: tracking actions as moving points with 3D CNNs
Shentong Mo, Jingfei Xia, Xiaoqing Tan +1
Spatio-temporal action recognition has been a challenging task that involves detecting where and when actions occur. Current state-of-the-art action detectors are mostly anchor-bas…
Self-Supervised 3D Face Reconstruction via Conditional Estimation
Yandong Wen, Weiyang Liu, Bhiksha Raj +1
We present a conditional estimation (CEST) framework to learn 3D facial parameters from 2D single-view images by self-supervised training from videos. CEST is based on the process…
The Right to Talk: An Audio-Visual Transformer Approach
Thanh-Dat Truong, Chi Nhan Duong, The De Vu +4
Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking)…
Controlled AutoEncoders to Generate Faces from Voices
Hao Liang, Lulan Yu, Guikang Xu +2
Multiple studies in the past have shown that there is a strong correlation between human vocal characteristics and facial features. However, existing approaches generate faces simp…