2 papers
cs.SD2022
Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast
Boqing Zhu, Kele Xu, Changjian Wang +4
We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination task…
cs.CV2022
S3E-GNN: Sparse Spatial Scene Embedding with Graph Neural Networks for Camera Relocalization
Ran Cheng, Xinyu Jiang, Yuan Chen +2
Camera relocalization is the key component of simultaneous localization and mapping (SLAM) systems. This paper proposes a learning-based approach, named Sparse Spatial Scene Embedd…