2 citations · 3 across the 2 of their papers we have counts for
3 papers
MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization
Alexander Kunitsyn, Maksim Kalashnikov, Maksim Dzabraev +1
In this work we present a new State-of-The-Art on the text-to-video retrieval task on MSR-VTT, LSMDC, MSVD, YouCook2 and TGIF obtained by a single model. Three different data sourc…
MDMMT: Multidomain Multimodal Transformer for Video Retrieval
Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov +1
We present a new state-of-the-art on the text to video retrieval task on MSRVTT and LSMDC benchmarks where our model outperforms all previous solutions by a large margin. Moreover,…
Mutual Modality Learning for Video Action Classification
Stepan Komkov, Maksim Dzabraev, Aleksandr Petiushko
The construction of models for video action classification progresses rapidly. However, the performance of those models can still be easily improved by ensembling with the same mod…