1 citations · 1 across the 15 of their papers we have counts for
16 papers
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
Ke Zhang, Yankang Liu, Roya Zandi +1
Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and…
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
Ruoxi Zhao, Maziar Raissi
Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual c…
VeraGrid-Agent: Tool-Augmented LLMs for Distribution Optimal Power Flow at the Grid Edge
Shivanshu Tripathi, Hamed Mohsenian-Rad, Maziar Raissi
Language models have demonstrated remarkable success in solving a wide range of tasks. However, answering complex scientific questions about the power flow often requires solving t…
Spectrogram-Based Joint Detection, Localization, and Classification of Events in Continuously Recorded IBR Waveforms
Shivanshu Tripathi, Maziar Raissi, Hamed Mohsenian-Rad
Continuously recorded high-resolution waveform measurements provide rich information about fast power system dynamics. However, they require automated methods to identify events. T…
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
Ke Zhang, Sahchit Chundur, Mohammad Javad Qomi +1
Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely…
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
Ke Zhang, Patricio Gallardo Candela, Sudhir Murthy +3
Lean verifies that a generated declaration is well typed, but not that it expresses the statement a user intended. We study two questions for autoformalization without canonical Le…