Pre-training with Graph Transformers
arXiv:2609.13844
Abstract
This article investigates pre-training strategies for graph transformers in the biochemistry domain. By conducting comprehensive experiments, the study reveals that supervised pre-training using computed properties as labels provides the highest performance gain on downstream tasks. The results also highlight the importance of constraining model capacity to mitigate overfitting in graph transformers.
4 pages, 1 table. DLG-KDD 2023 workshop paper