1 paper
Arkaprava Sinha, Dominick Reilly, Francois Bremond +2
The introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, t…