1 paper
Guozhen Zhang, Jingyu Liu, Shengming Cao +4
Recently, the remarkable success of pre-trained Vision Transformers (ViTs) from image-text matching has sparked an interest in image-to-video adaptation. However, most current appr…