1 paper
Viraat Aryabumi, Yixuan Su, Raymond Ma +6
Including code in the pre-training data mixture, even for models not specifically designed for code, has become a common practice in LLMs pre-training. While there has been anecdot…