10 papers
Divergence Decoding: Training-Free Capability Fusion
Yimi Wang, Hao Li, Shuo Yang +6
The paper proposes Divergence Decoding, a training‑free method that dynamically routes token generation between a generalist LLM and a domain‑specialist LLM using Jensen‑Shannon di…
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
Chengchun Liu, Zhiyuan Yan, Li Yuan +5
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are spar…
From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models
Hongyu Guo, Hao Li, He Cao +2
Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may o…
Agentic reinforcement learning empowers next-generation chemical language models for molecular design and synthesis
Hao Li, He Cao, Shenyao Peng +7
Language models are revolutionizing the biochemistry domain, assisting scientists in drug design and chemical synthesis with high efficiency. Yet current approaches struggle betwee…
Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations
Hao Li, He Cao, Bin Feng +6
While large language models (LLMs) with Chain-of-Thought (CoT) reasoning excel in mathematics and coding, their potential for systematic reasoning in chemistry, a domain demanding…
Rethinking Text-based Protein Understanding: Retrieval or LLM?
Juntong Wu, Zijing Liu, He Cao +6
In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding. Current approaches focus on integrating protein-…