3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2024
SnAG: Scalable and Accurate Video Grounding
Fangzhou Mu, Sicheng Mo, Yin Li
Temporal grounding of text descriptions in videos is a central problem in vision-language learning and video understanding. Existing methods often prioritize accuracy over scalabil…
cs.CV2023★ 3 cited
A Review of Adversarial Attacks in Computer Vision
Yutong Zhang, Yao Li, Yin Li +1
Deep neural networks have been widely used in various downstream tasks, especially those safety-critical scenario such as autonomous driving, but deep networks are often threatened…