most citedCAT: Cross Attention in Vision Transformer

17 citations · 18 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV2024

E4C: Enhance Editability for Text-Based Image Editing by Harnessing Efficient CLIP Guidance

Tianrui Huang, Pu Cao, Lu Yang +4

Diffusion-based image editing is a composite process of preserving the source image content and generating new content or applying modifications. While current editing approaches h…

cs.CV2023

CoT-MISR:Marrying Convolution and Transformer for Multi-Image Super-Resolution

Mingming Xiu, Yang Nie, Qing Song +1

As a method of image restoration, image super-resolution has been extensively studied at first. How to transform a low-resolution image to restore its high-resolution image informa…

cs.CV2023

Faster Learning of Temporal Action Proposal via Sparse Multilevel Boundary Generator

Qing Song, Yang Zhou, Mengjie Hu +1

Temporal action localization in videos presents significant challenges in the field of computer vision. While the boundary-sensitive method has been widely adopted, its limitations…

cs.CV20221 cited

SGM-Net: Semantic Guided Matting Net

Qing Song, Wenfeng Sun, Donghan Yang +2

Human matting refers to extracting human parts from natural images with high quality, including human detail information such as hair, glasses, hat, etc. This technology plays an e…

cs.CV202117 cited

CAT: Cross Attention in Vision Transformer

Hezheng Lin, Xing Cheng, Xiangyu Wu +5

Since Transformer has found widespread use in NLP, the potential of Transformer in CV has been realized and has inspired many new approaches. However, the computation required for…