Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition
Jinnuo Liu, Yue Peng, Jinhan Niu +1
Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a function name: models must coordin…
cs.AI2025
StepFun-Prover Preview: Let's Think and Verify Step by Step
Shijie Shang, Ruosi Wan, Yue Peng +4
We present StepFun-Prover Preview, a large language model designed for formal theorem proving through tool-integrated reasoning. Using a reinforcement learning pipeline that incorp…