1 paper · 1 filter
Hyo Jin Jon, Longbin Jin, Eun Yi Kim
CLIP has demonstrated strong generalization in visual domains through natural language supervision, even for video action recognition. However, most existing approaches that adapt…