1 paper
Hu Xu, Gargi Ghosh, Po-Yao Huang +5
We present a simplified, task-agnostic multi-modal pre-training approach that can accept either video or text input, or both for a variety of end tasks. Existing pre-training are t…