2 papers
cs.CL2026
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
Yuru Wang, Lejun Cheng, Yuxin Zuo +14
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move b…
cs.CL2025
DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base
Song Mao, Lejun Cheng, Pinlong Cai +3
Large Language Models (LLMs) have demonstrated remarkable capabilities in various applications. However, their use as writing assistants in specialized domains like finance, medici…