6 citations · 6 across the 1 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
Zheyu Zhang, Ziqi Pang, Shixing Chen +3
Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens…
cs.CV2022★ 6 cited
Robust Cross-Modal Representation Learning with Progressive Self-Distillation
Alex Andonian, Shixing Chen, Raffay Hamid
The learning objective of vision-language approach of CLIP does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets,…
cs.CV2021
Shot Contrastive Self-Supervised Learning for Scene Boundary Detection
Shixing Chen, Xiaohan Nie, David Fan +3
Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boun…