2 citations · 2 across the 3 of their papers we have counts for
5 papers
Cross-Modal Discrete Representation Learning
Alexander H. Liu, SouYoung Jin, Cheng-I Jeff Lai +3
Recent advances in representation learning have demonstrated an ability to represent information from different modalities such as video, text, and audio in a single high-level emb…
Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions
Mathew Monfort, SouYoung Jin, Alexander Liu +4
When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information…
Automatic adaptation of object detectors to new domains using self-training
Aruni RoyChowdhury, Prithvijit Chakrabarty, Ashish Singh +4
This work addresses the unsupervised adaptation of an existing object detector to a new target domain. We assume that a large number of unlabeled videos from this domain are readil…
Unsupervised Hard Example Mining from Videos for Improved Object Detection
SouYoung Jin, Aruni RoyChowdhury, Huaizu Jiang +4
Important gains have recently been obtained in object detection by using training objectives that focus on {\em hard negative} examples, i.e., negative examples that are currently…
End-to-end Face Detection and Cast Grouping in Movies Using Erdős-Rényi Clustering
SouYoung Jin, Hang Su, Chris Stauffer +1
We present an end-to-end system for detecting and clustering faces by identity in full-length movies. Unlike works that start with a predefined set of detected faces, we consider t…