4 papers · 1 filter
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
Ke Zhang, Yankang Liu, Roya Zandi +1
Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and…
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
Ke Zhang, Sahchit Chundur, Mohammad Javad Qomi +1
Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely…
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
Ke Zhang, Patricio Gallardo Candela, Sudhir Murthy +3
Lean verifies that a generated declaration is well typed, but not that it expresses the statement a user intended. We study two questions for autoformalization without canonical Le…
Pioneer Agent: Continual Improvement of Small Language Models in Production
Dhruv Atreja, Julia White, Nikhil Nayak +5
Small language models are attractive for production deployment due to their low cost, fast inference, and ease of specialization. However, adapting them to a specific task remains…