132 citations · 345 across the 20 of their papers we have counts for
5 papers · 1 filter
First Place Solution of KDD Cup 2021 & OGB Large-Scale Challenge Graph Prediction Track
Chengxuan Ying, Mingqi Yang, Shuxin Zheng +7
In this technical report, we present our solution of KDD Cup 2021 OGB Large-Scale Challenge - PCQM4M-LSC Track. We adopt Graphormer and ExpC as our basic models. We train each mode…
Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
Shengjie Luo, Shanda Li, Tianle Cai +6
The attention module, which is a crucial component in Transformer, cannot scale efficiently to long sequences due to its quadratic complexity. Many works focus on approximating the…
Do Transformers Really Perform Bad for Graph Representation?
Chengxuan Ying, Tianle Cai, Shengjie Luo +5
The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision. Yet, it has not achieved competitive performance…
How could Neural Networks understand Programs?
Dinglan Peng, Shuxin Zheng, Yatao Li +3
Semantic understanding of programs is a fundamental problem for programming language processing (PLP). Recent works that learn representations of code based on pre-training techniq…
Revisiting Language Encoding in Learning Multilingual Representations
Shengjie Luo, Kaiyuan Gao, Shuxin Zheng +4
Transformer has demonstrated its great power to learn contextual word representations for multiple languages in a single model. To process multilingual sentences in the model, a le…