1 paper
John Pougue Biyong, Bo Wang, Terry Lyons +1
Relying on large pretrained language models such as Bidirectional Encoder Representations from Transformers (BERT) for encoding and adding a simple prediction layer has led to impr…