Developing ChemDFM as a large language foundation model for chemistry
arXiv:2401.14818 · doi:10.1016/j.xcrp.2025.102523
Abstract
Artificial intelligence (AI) has played an increasingly important role in chemical research. However, most models currently used in chemistry are specialist models that require training and tuning for specific tasks. A more generic and efficient solution would be an AI model that could address many tasks and support free-form dialogue in the broad field of chemistry. In its utmost form, such a generalist AI chemist could be referred to as Chemical General Intelligence. Large language models (LLMs) have recently logged tremendous success in the general domain of natural language processing, showing emerging task generalization and free-form dialogue capabilities. However, domain knowledge of chemistry is largely missing when training general-domain LLMs. The lack of such knowledge greatly hinders the performance of generalist LLMs in the field of chemistry. To this end, we develop ChemDFM, a pioneering LLM for chemistry trained on 34B tokens from chemical literature and textbooks, and fine-tuned using 2.7M instructions. As a result, it can understand and reason with chemical knowledge in free-form dialogue. Quantitative evaluations show that ChemDFM significantly surpasses most representative open-source LLMs. It outperforms GPT-4 on a great portion of chemical tasks, despite the substantial size difference. We have open-sourced the inference codes, evaluation datasets, and model weights of ChemDFM on Huggingface (https://huggingface.co/OpenDFM/ChemDFM-v1.0-13B).
10 pages, 12 figures, 12 tables. Published on Cell Report Physical Science, DOI: https://doi.org/10.1016/j.xcrp.2025.102523
References in corpus (24)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Large Language Models are Zero-Shot Reasoners
- Molecular Transformer - A Model for Uncertainty-Calibrated Chemical Reaction Prediction
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Galactica: A Large Language Model for Science
- Root-aligned SMILES: A Tight Representation for Chemical Reaction Prediction
- What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- as a Two-Step Graph Generative Models for Retrosynthesis Prediction
- EduChat: A Large-Scale Language Model-based Chatbot System for Intelligent Education
- DARWIN Series: Domain Specific Large Language Models for Natural Science
- Unifying Molecular and Textual Representations via Multi-task Language Modelling
- A Closer Look at AUROC and AUPRC under Class Imbalance
- Are large language models superhuman chemists?
- InstructMol: Multi-Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery
- Multi-Modal Representation Learning for Molecular Property Prediction: Sequence, Graph, Geometry
- A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
- 3D Molecular Generation via Virtual Dynamics
- UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation