5 papers
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
Sihan Hu, Lyuhan Huang, Youjin Deng +1
SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as w…
Emergent Slow Thinking in LLMs as Inverse Tree Freezing
Sihan Hu, Xiansheng Cai, Yuan Huang +5
Reinforcement learning with verifiable rewards (RLVR) enables large language models to acquire slow, multi-step reasoning from sparse final-answer signals. We provide a statistical…
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
Zichen Wen, Boxue Yang, Shuang Chen +30
We present Innovator-VL, a scientific multimodal large language model designed to advance understanding and reasoning across diverse scientific domains while maintaining excellent…
Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
Yu Li, Yuan Huang, Tao Wang +19
Most scientific materials compress reasoning, presenting conclusions while omitting the derivational chains that justify them. This compression hinders verification by lacking expl…
Learning-at-Criticality in Large Language Models for Quantum Field Theory and Beyond
Xiansheng Cai, Sihan Hu, Tao Wang +4
Fundamental physics often confronts complex symbolic problems with few guiding exemplars or established principles. While artificial intelligence (AI) offers promise, its typical n…