16 citations · 24 across the 4 of their papers we have counts for
5 papers
Efficient Modulation for Vision Networks
Xu Ma, Xiyang Dai, Jianwei Yang +4
In this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context…
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Bin Xiao, Haiping Wu, Weijian Xu +6
We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks. While existing larg…
LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction Following
Cheng-Fu Yang, Yen-Chun Chen, Jianwei Yang +4
End-to-end Transformers have demonstrated an impressive success rate for Embodied Instruction Following when the environment has been seen in training. However, they tend to strugg…
Learning from Rich Semantics and Coarse Locations for Long-tailed Object Detection
Lingchen Meng, Xiyang Dai, Jianwei Yang +7
Long-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to expl…
Exploring Invariance in Images through One-way Wave Equations
Yinpeng Chen, Dongdong Chen, Xiyang Dai +5
In this paper, we empirically reveal an invariance over images-images share a set of one-way wave equations with latent speeds. Each image is uniquely associated with a solution to…