7 citations · 10 across the 4 of their papers we have counts for
4 papers
An Empirical Study of Graphormer on Large-Scale Molecular Modeling Datasets
Yu Shi, Shuxin Zheng, Guolin Ke +7
This technical note describes the recent updates of Graphormer, including architecture design modifications, and the adaption to 3D molecular dynamics simulation. The "Graphormer-V…
First Place Solution of KDD Cup 2021 & OGB Large-Scale Challenge Graph Prediction Track
Chengxuan Ying, Mingqi Yang, Shuxin Zheng +7
In this technical report, we present our solution of KDD Cup 2021 OGB Large-Scale Challenge - PCQM4M-LSC Track. We adopt Graphormer and ExpC as our basic models. We train each mode…
Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
Shengjie Luo, Shanda Li, Tianle Cai +6
The attention module, which is a crucial component in Transformer, cannot scale efficiently to long sequences due to its quadratic complexity. Many works focus on approximating the…
Revisiting Language Encoding in Learning Multilingual Representations
Shengjie Luo, Kaiyuan Gao, Shuxin Zheng +4
Transformer has demonstrated its great power to learn contextual word representations for multiple languages in a single model. To process multilingual sentences in the model, a le…