most citedLite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR

8 citations · 25 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20234 cited

A Strong and Reproducible Object Detector with Only Public Datasets

Tianhe Ren, Jianwei Yang, Shilong Liu +6

This work presents Focal-Stable-DINO, a strong and reproducible object detection model which achieves 64.6 AP on COCO val2017 and 64.8 AP on COCO test-dev using only 700M parameter…

cs.CV20231 cited

DisCo-CLIP: A Distributed Contrastive Loss for Memory Efficient CLIP Training

Yihao Chen, Xianbiao Qi, Jianan Wang +1

We propose DisCo-CLIP, a distributed memory-efficient CLIP training approach, to reduce the memory consumption of contrastive loss when training contrastive learning models. Our ap…

cs.CV20235 cited

Detection Transformer with Stable Matching

Shilong Liu, Tianhe Ren, Jiayu Chen +8

This paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is cau…

cs.CV20237 cited

MP-Former: Mask-Piloted Transformer for Image Segmentation

Hao Zhang, Feng Li, Huaizhe Xu +4

We present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers fr…

cs.CV20238 cited

Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR

Feng Li, Ailing Zeng, Shilong Liu +4

Recent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in t…

cs.CV2023

Introducing Depth into Transformer-based 3D Object Detection

Hao Zhang, Hongyang Li, Ailing Zeng +4

In this paper, we present DAT, a Depth-Aware Transformer framework designed for camera-based 3D detection. Our model is based on observing two major issues in existing methods: lar…