5 papers
An Empirical Study on Clustering Pretrained Embeddings: Is Deep Strictly Better?
Tyler R. Scott, Ting Liu, Michael C. Mozer +1
Recent research in clustering face embeddings has found that unsupervised, shallow, heuristic-based methods -- including -means and hierarchical agglomerative clustering -- unde…
AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch +8
Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, a…
Modeling Uncertainty with Hedged Instance Embedding
Seong Joon Oh, Kevin Murphy, Jiyan Pan +3
Instance embeddings are an efficient and versatile image representation that facilitates applications like recognition, verification, retrieval, and clustering. Many metric learnin…
AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies
Sourish Chaudhuri, Joseph Roth, Daniel P. W. Ellis +8
Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio-…
Finding your Lookalike: Measuring Face Similarity Rather than Face Identity
Amir Sadovnik, Wassim Gharbi, Thanh Vu +1
Face images are one of the main areas of focus for computer vision, receiving on a wide variety of tasks. Although face recognition is probably the most widely researched, many oth…