21 citations · 21 across the 3 of their papers we have counts for
4 papers · 1 filter
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
Sihan Chen, Handong Li, Qunbo Wang +4
Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient a…
OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
Jing Liu, Xinxin Zhu, Fei Liu +8
In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…
Learning Nonparametric Human Mesh Reconstruction from a Single Image without Ground Truth Meshes
Kevin Lin, Lijuan Wang, Ying Jin +2
Nonparametric approaches have shown promising results on reconstructing 3D human mesh from a single monocular image. Unlike previous approaches that use a parametric human model li…
Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning
Dianqi Li, Qiuyuan Huang, Xiaodong He +2
We study how to generate captions that are not only accurate in describing an image but also discriminative across different images. The problem is both fundamental and interesting…