16 citations · 26 across the 2 of their papers we have counts for
8 papers
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Bin Xiao, Haiping Wu, Weijian Xu +6
We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks. While existing larg…
MM-VID: Advancing Video Understanding with GPT-4V(ision)
Kevin Lin, Faisal Ahmed, Linjie Li +9
We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…
Florence: A New Foundation Model for Computer Vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen +20
Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human visio…
UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
Jianfeng Wang, Xiaowei Hu, Zhe Gan +5
In this paper, we propose a single UniFied transfOrmer (UFO), which is capable of processing either unimodal inputs (e.g., image or language) or multimodal inputs (e.g., the concat…
SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning
Kevin Lin, Linjie Li, Chung-Ching Lin +5
The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on vid…
Scaling Up Vision-Language Pre-training for Image Captioning
Xiaowei Hu, Zhe Gan, Jianfeng Wang +4
In recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important fact…