most citedAudio-Visual Transformer Based Crowd Counting

2 citations · 6 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV20222 cited

Explicitly Increasing Input Information Density for Vision Transformers on Small Datasets

Xiangyu Chen, Ying Qin, Wenju Xu +3

Vision Transformers have attracted a lot of attention recently since the successful implementation of Vision Transformer (ViT) on vision tasks. With vision Transformers, specifical…

cs.CV20222 cited

Accumulated Trivial Attention Matters in Vision Transformers on Small Datasets

Xiangyu Chen, Qinghao Hu, Kaidong Li +2

Vision Transformers has demonstrated competitive performance on computer vision tasks benefiting from their ability to capture long-range dependencies with multi-head self-attentio…

cs.CV2022

Dilated Continuous Random Field for Semantic Segmentation

Xi Mo, Xiangyu Chen, Cuncong Zhong +3

Mean field approximation methodology has laid the foundation of modern Continuous Random Field (CRF) based solutions for the refinement of semantic segmentation. In this paper, we…

cs.CV20212 cited

Audio-Visual Transformer Based Crowd Counting

Usman Sajid, Xiangyu Chen, Hasan Sajid +2

Crowd estimation is a very challenging problem. The most recent study tries to exploit auditory information to aid the visual models, however, the performance is limited due to the…

cs.CV2021

Few-Shot Learning by Integrating Spatial and Frequency Representation

Xiangyu Chen, Guanghui Wang

Human beings can recognize new objects with only a few labeled examples, however, few-shot learning remains a challenging problem for machine learning systems. Most previous algori…