Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
Yuyang Wu, Yue Huang, Shuaike Shen +8
Large Language Models (LLMs) have become increasingly capable as tool-using agents, with benchmarks spanning diverse general agentic tasks. Yet rigorous evaluation of scientific to…
cs.AI2026
SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources
Shuaike Shen, Wenduo Cheng, Mingqian Ma +3
Modern scientific ecosystems are rich in procedural knowledge across repositories, APIs, scripts, notebooks, documentation, databases, and papers, yet much of this knowledge remain…