4 papers · 1 filter
Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
Xiaoyu Zhu, Hao Zhou, Pengfei Xing +6
In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel…
An Empirical Study on Clustering Pretrained Embeddings: Is Deep Strictly Better?
Tyler R. Scott, Ting Liu, Michael C. Mozer +1
Recent research in clustering face embeddings has found that unsupervised, shallow, heuristic-based methods -- including -means and hierarchical agglomerative clustering -- unde…
AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch +8
Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, a…
Finding your Lookalike: Measuring Face Similarity Rather than Face Identity
Amir Sadovnik, Wassim Gharbi, Thanh Vu +1
Face images are one of the main areas of focus for computer vision, receiving on a wide variety of tasks. Although face recognition is probably the most widely researched, many oth…