24 citations · 30 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 24 cited
MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining
Pengyuan Lyu, Chengquan Zhang, Shanshan Liu +7
Text images contain both visual and linguistic information. However, existing pre-training techniques for text recognition mainly focus on either visual representation learning or…
cs.CV2022★ 6 cited
Image Captioning In the Transformer Age
Yang Xu, Li Li, Haiyang Xu +3
Image Captioning (IC) has achieved astonishing developments by incorporating various techniques into the CNN-RNN encoder-decoder architecture. However, since CNN and RNN do not sha…