papers

Publications (13)

cs.CL2024

ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity

Tong Xie, Hanzhi Zhang, Shaozhou Wang +5

Natural Language Processing (NLP) is widely used to supply summarization ability from long context to structured information. However, extracting structured knowledge from scientif…

cs.CL2023

Large Language Models as Master Key: Unlocking the Secrets of Materials Science with GPT

Tong Xie, Yuwei Wan, Wei Huang +8

The amount of data has growing significance in exploring cutting-edge materials and a number of datasets have been generated either by hand or automated approaches. However, the ma…

cs.CL2024

SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation

Yuwei Wan, Yixuan Liu, Aswathy Ajith +6

We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language…

cs.CL2025

DARWIN 1.5: Large Language Models as Materials Science Adapted Learners

Tong Xie, Yuwei Wan, Yixuan Liu +8

Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as hig…

cs.LG2026

MiST: Understanding the Role of Mid-Stage Scientific Training in Developing Chemical Reasoning Models

Andres M Bran, Tong Xie, Shai Pranesh +9

Large Language Models can develop reasoning capabilities through online fine-tuning with rule-based rewards. However, recent studies reveal a critical constraint: reinforcement lea…

cs.AI2025

DrugMCTS: a drug repurposing framework combining multi-agent, RAG and Monte Carlo Tree Search

Zerui Yang, Yuwei Wan, Siyu Yan +4

Recent advances in large language models have demonstrated considerable potential in scientific domains such as drug repositioning. However, their effectiveness remains constrained…