activity
20182022
most citedERNIE-Search: Bridging Cross-Encoder with Dual-Encoder via Self On-the-fly Distillation for Dense Passage Retrieval

19 citations · 55 across the 8 of their papers we have counts for

collaborators

12 papers

cs.CL2022

DialogConv: A Lightweight Fully Convolutional Network for Multi-view Response Selection

Yongkang Liu, Shi Feng, Wei Gao +2

Current end-to-end retrieval-based dialogue systems are mainly based on Recurrent Neural Networks or Transformers with attention mechanisms. Although promising results have been ac…

cs.CL20226 cited

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Qiming Peng, Yinxu Pan, Wenjin Wang +12

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and u…

cs.CV20223 cited

ERNIE-mmLayout: Multi-grained MultiModal Transformer for Document Understanding

Wenjin Wang, Zhengjie Huang, Bin Luo +8

Recent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approa…

cs.CL202219 cited

ERNIE-Search: Bridging Cross-Encoder with Dual-Encoder via Self On-the-fly Distillation for Dense Passage Retrieval

Yuxiang Lu, Yiding Liu, Jiaxiang Liu +8

Neural retrievers based on pre-trained language models (PLMs), such as dual-encoders, have achieved promising performance on the task of open-domain question answering (QA). Their…

cs.CL20223 cited

Simple and Effective Relation-based Embedding Propagation for Knowledge Representation Learning

Huijuan Wang, Siming Dai, Weiyue Su +7

Relational graph neural networks have garnered particular attention to encode graph context in knowledge graphs (KGs). Although they achieved competitive performance on small KGs,…

cs.CL20226 cited

ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention

Yang Liu, Jiaxiang Liu, Li Chen +7

Sparse Transformer has recently attracted a lot of attention since the ability for reducing the quadratic dependency on the sequence length. We argue that two factors, information…