21 citations · 22 across the 2 of their papers we have counts for
1 paper · 1 filter
Jing Liu, Xinxin Zhu, Fei Liu +8
In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…