5 papers
A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
Feiyang Cai, Guijuan He, Yi Hu +7
Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential for enabling large language models (LLMs…
MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation
Feiyang Cai, Jiahui Bai, Tao Tang +7
Precise recognition, editing, and generation of molecules are essential prerequisites for both chemists and AI systems tackling various chemical tasks. We present MolLangBench, a c…
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
Kunning Li, Jianbin Guo, Zhaoyang Shang +5
The emergence of Large Language Models (LLMs) within the Traditional Chinese Medicine (TCM) domain presents an urgent need to assess their clinical application capabilities. Howeve…
ChemFM as a Scaling Law Guided Foundation Model Pre-trained on Informative Chemicals
Feiyang Cai, Katelin Zacour, Tianyu Zhu +6
Traditional AI methods often rely on task-specific model designs and training, which constrain both the scalability of model size and generalization across different tasks. Here, w…
From Intention To Implementation: Automating Biomedical Research via LLMs
Yi Luo, Linghang Shi, Yihao Li +4
Conventional biomedical research is increasingly labor-intensive due to the exponential growth of scientific literature and datasets. Artificial intelligence (AI), particularly Lar…