activity
20212023
most citedFlorence-2: Advancing a Unified Representation for a Variety of Vision Tasks

16 citations · 50 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV202316 cited

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Bin Xiao, Haiping Wu, Weijian Xu +6

We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks. While existing larg…

cs.CL20231 cited

LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction Following

Cheng-Fu Yang, Yen-Chun Chen, Jianwei Yang +4

End-to-end Transformers have demonstrated an impressive success rate for Embodied Instruction Following when the environment has been seen in training. However, they tend to strugg…

cs.CV20234 cited

Learning from Rich Semantics and Coarse Locations for Long-tailed Object Detection

Lingchen Meng, Xiyang Dai, Jianwei Yang +7

Long-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to expl…

cs.CV20237 cited

ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Junke Wang, Dongdong Chen, Chong Luo +4

Existing deep video models are limited by specific tasks, fixed input-output spaces, and poor generalization capabilities, making it difficult to deploy them in real-world scenario…

cs.CV20235 cited

Layer Grafted Pre-training: Bridging Contrastive Learning And Masked Image Modeling For Label-Efficient Representations

Ziyu Jiang, Yinpeng Chen, Mengchen Liu +5

Recently, both Contrastive Learning (CL) and Mask Image Modeling (MIM) demonstrate that self-supervision is powerful to learn good representations. However, naively combining them…

cs.CV20224 cited

Video Mobile-Former: Video Recognition with Efficient Global Spatial-temporal Modeling

Rui Wang, Zuxuan Wu, Dongdong Chen +6

Transformer-based models have achieved top performance on major video recognition benchmarks. Benefiting from the self-attention mechanism, these models show stronger ability of mo…