15 citations · 42 across the 5 of their papers we have counts for
6 papers
MoCoViT: Mobile Convolutional Vision Transformer
Hailong Ma, Xin Xia, Xing Wang +3
Recently, Transformer networks have achieved impressive results on a variety of vision tasks. However, most of them are computationally expensive and not suitable for real-world mo…
Dressing in the Wild by Watching Dance Videos
Xin Dong, Fuwei Zhao, Zhenyu Xie +6
While significant progress has been made in garment transfer, one of the most applicable directions of human-centric image generation, existing works overlook the in-the-wild image…
Fast Convergence of DETR with Spatially Modulated Co-Attention
Peng Gao, Minghang Zheng, Xiaogang Wang +2
The recently proposed Detection Transformer (DETR) model successfully applies Transformer to objects detection and achieves comparable performance with two-stage object detection f…
Fast Convergence of DETR with Spatially Modulated Co-Attention
Peng Gao, Minghang Zheng, Xiaogang Wang +2
The recently proposed Detection Transformer (DETR) model successfully applies Transformer to objects detection and achieves comparable performance with two-stage object detection f…
End-to-End Object Detection with Adaptive Clustering Transformer
Minghang Zheng, Peng Gao, Renrui Zhang +4
End-to-end Object Detection with Transformer (DETR)proposes to perform object detection with Transformer and achieve comparable performance with two-stage object detection like Fas…
Identify Equivalent Frames
Xuemei Chen, Yang Chu, Min Zheng
A frame is an overcomplete set that can represent vectors(signals) faithfully and stably. Two frames are equivalent if signals can be essentially represented in the same way, which…