bioinformatics 1biological corpora 1evaluation benchmarks 1large language models 1pretraining data 1
From the 1 of 3 linked papers with an AI index.
3 papers
q-bio.QM2026
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
Hyunjin Seo, Hyeon Hwang, Gyubok Lee +7
The paper introduces TheBioCollection, a 52.6‑billion‑token unified corpus that aggregates diverse biological resources for pre‑training large language models, and shows that train…
cs.LG2026
Predicting LLM Reasoning Performance with Small Proxy Model
Woosung Koh, Juyoung Suk, Sungjun Han +2
Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize datasets before scaling up. However, this approach be…
cs.CL2025
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
Seyoung Song, Seogyeong Jeong, Eunsu Kim +4
Evaluating text generation capabilities of large language models (LLMs) is challenging, particularly for low-resource languages where methods for direct assessment are scarce. We p…