32 citations · 43 across the 3 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2022★ 4 cited
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers
Zhewei Yao, Xiaoxia Wu, Conglong Li +4
Large-scale transformer models have become the de-facto architectures for various machine learning applications, e.g., CV and NLP. However, those large models also introduce prohib…
cs.CL2021★ 7 cited
LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
Te-Lin Wu, Cheng Li, Mingyang Zhang +3
Document layout comprises both structural and visual (eg. font-sizes) information that is vital but often ignored by machine learning models. The few existing models which do use l…
cs.CL2020★ 32 cited
Downstream Model Design of Pre-trained Language Model for Relation Extraction Task
Cheng Li, Ye Tian
Supervised relation extraction methods based on deep neural network play an important role in the recent information extraction field. However, at present, their performance still…