4 citations · 4 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023
Language-aware Multiple Datasets Detection Pretraining for DETRs
Jing Hao, Song Chen, Xiaodi Wang +1
Pretraining on large-scale datasets can boost the performance of object detectors while the annotated datasets for object detection are hard to scale up due to the high labor cost.…
cs.CV2022★ 4 cited
MAFormer: A Transformer Network with Multi-scale Attention Fusion for Visual Recognition
Yunhao Wang, Huixin Sun, Xiaodi Wang +6
Vision Transformer and its variants have demonstrated great potential in various computer vision tasks. But conventional vision transformers often focus on global dependency at a c…