2 papers
cs.AI2026
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Songlin Bai, Xintong Wang, Linlin Yu +12
In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every parameter must respect a regula…
cs.AI2025
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
Rui Min, Zile Qiao, Ze Xu +18
Foundation agents have rapidly advanced in their ability to reason and interact with real environments, making the evaluation of their core capabilities increasingly important. Whi…