q-bio.QM2026
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
Hyunjin Seo, Hyeon Hwang, Gyubok Lee +7
The paper introduces TheBioCollection, a 52.6‑billion‑token unified corpus that aggregates diverse biological resources for pre‑training large language models, and shows that train…
#large language models#biological corpora#pretraining data#bioinformatics