3 papers
cs.CL2026
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
Jiaming Wang, Zhe Tang, Zehao Jin +5
As large language models (LLMs) are widely deployed as domain-specific agents, many benchmarks have been proposed to evaluate their ability to follow instructions and make decision…
cs.CL2025
Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
Jiaming wang, Yunke Zhao, Peng Ding +8
The capability to precisely adhere to instructions is a cornerstone for Large Language Models (LLMs) to function as dependable agents in real-world scenarios. However, confronted w…
cs.AI2025
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
Xiaoyang Chen, Xinan Dai, Yu Du +28
To advance the mathematical proficiency of large language models (LLMs), the DeepMath team has launched an open-source initiative aimed at developing an open mathematical LLM and s…