1 paper
Song Jin, Zhongtao Jiang, Chenglei Shen +5
Large-scale video retrieval requires embedding models to encode long and diverse videos under tight visual-input and inference budgets. Existing methods typically sample a small, f…