8 citations · 10 across the 2 of their papers we have counts for
4 papers
CodeBPE: Investigating Subtokenization Options for Large Language Model Pretraining on Source Code
Nadezhda Chirkova, Sergey Troshin
Recent works have widely adopted large language model pretraining for source code, suggested source code-specific pretraining objectives and investigated the applicability of vario…
Probing Pretrained Models of Source Code
Sergey Troshin, Nadezhda Chirkova
Deep learning models are widely used for solving challenging code processing tasks, such as code generation or code summarization. Traditionally, a specific model architecture was…
A Simple Approach for Handling Out-of-Vocabulary Identifiers in Deep Learning for Source Code
Nadezhda Chirkova, Sergey Troshin
There is an emerging interest in the application of natural language processing models to source code processing tasks. One of the major problems in applying deep learning to softw…
Empirical Study of Transformers for Source Code
Nadezhda Chirkova, Sergey Troshin
Initially developed for natural language processing (NLP), Transformers are now widely used for source code processing, due to the format similarity between source code and text. I…