most citedFlorence-2: Advancing a Unified Representation for a Variety of Vision Tasks

16 citations · 26 across the 2 of their papers we have counts for

collaborators

8 papers

cs.CV2023★ 16 cited

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Bin Xiao, Haiping Wu, Weijian Xu +6

We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks. While existing larg…

cs.CV2023★ 10 cited

MM-VID: Advancing Video Understanding with GPT-4V(ision)

Kevin Lin, Faisal Ahmed, Linjie Li +9

We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…

cs.CV2021

Florence: A New Foundation Model for Computer Vision

Lu Yuan, Dongdong Chen, Yi-Ling Chen +20

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human visio…

cs.CV2021

UFO: A UniFied TransfOrmer for Vision-Language Representation Learning

Jianfeng Wang, Xiaowei Hu, Zhe Gan +5

In this paper, we propose a single UniFied transfOrmer (UFO), which is capable of processing either unimodal inputs (e.g., image or language) or multimodal inputs (e.g., the concat…

cs.CV2021

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

Kevin Lin, Linjie Li, Chung-Ching Lin +5

The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on vid…

cs.CV2021

Scaling Up Vision-Language Pre-training for Image Captioning

Xiaowei Hu, Zhe Gan, Jianfeng Wang +4

In recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important fact…