2 papers
cs.SE2026
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
Sihan Hu, Lyuhan Huang, Youjin Deng +1
SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as w…
cs.AI2026
Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
Yu Li, Yuan Huang, Tao Wang +19
Most scientific materials compress reasoning, presenting conclusions while omitting the derivational chains that justify them. This compression hinders verification by lacking expl…