128 citations · 135 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 4 cited
Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning
Weicong Liang, Yuhui Yuan, Henghui Ding +6
Vision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. M…
cs.CV2021★ 128 cited
HRFormer: High-Resolution Transformer for Dense Prediction
Yuhui Yuan, Rao Fu, Lang Huang +4
We present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that prod…
cs.CL2021★ 3 cited
ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents
Weihong Lin, Qifang Gao, Lei Sun +4
Recent grid-based document representations like BERTgrid allow the simultaneous encoding of the textual and layout information of a document in a 2D feature map so that state-of-th…