6 citations · 6 across the 1 of their papers we have counts for
2 papers
cs.CV2022★ 6 cited
Robust Cross-Modal Representation Learning with Progressive Self-Distillation
Alex Andonian, Shixing Chen, Raffay Hamid
The learning objective of vision-language approach of CLIP does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets,…
cs.CV2021
Shot Contrastive Self-Supervised Learning for Scene Boundary Detection
Shixing Chen, Xiaohan Nie, David Fan +3
Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boun…