6 citations · 8 across the 2 of their papers we have counts for
3 papers
cs.CV2022★ 6 cited
Robust Cross-Modal Representation Learning with Progressive Self-Distillation
Alex Andonian, Shixing Chen, Raffay Hamid
The learning objective of vision-language approach of CLIP does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets,…
cs.CV2022★ 2 cited
Depth-Guided Sparse Structure-from-Motion for Movies and TV Shows
Sheng Liu, Xiaohan Nie, Raffay Hamid
Existing approaches for Structure from Motion (SfM) produce impressive 3-D reconstruction results especially when using imagery captured with large parallax. However, to create eng…
cs.CV2021
Shot Contrastive Self-Supervised Learning for Scene Boundary Detection
Shixing Chen, Xiaohan Nie, David Fan +3
Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boun…