8 papers
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
Yizhou Wang, Chen Tang, Han Deng +29
We present a scientific reasoning foundation model that aligns natural language with heterogeneous scientific representations. The model is pretrained on a 206B-token corpus spanni…
GoRA: Gradient-driven Adaptive Low Rank Adaptation
Haonan He, Peng Ye, Yuchen Ren +4
Low-Rank Adaptation (LoRA) is a crucial method for efficiently fine-tuning large language models (LLMs), with its effectiveness influenced by two key factors: rank selection and we…
A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
Ming Hu, Chenglong Ma, Wei Li +117
Scientific Large Language Models (Sci-LLMs) are transforming how knowledge is represented, integrated, and applied in scientific research, yet their progress is shaped by the compl…
Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models
Haonan He, Yuchen Ren, Yining Tang +12
Large language models (LLMs) have shown remarkable capabilities in general domains, but their application to multi-omics biology remains underexplored. To address this gap, we intr…
Intern-S1: A Scientific Multimodal Foundation Model
Lei Bai, Zhongrui Cai, Yuhang Cao +173
In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that…
Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA
Lifeng Qiao, Peng Ye, Yuchen Ren +5
Foundation models have made significant strides in understanding the genomic language of DNA sequences. However, previous models typically adopt the tokenization methods designed f…