2 papers
cs.AI2026
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Songlin Bai, Xintong Wang, Linlin Yu +12
In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every parameter must respect a regula…
cs.CL2025
ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions
Xingwei He, Qianru Zhang, Pengfei Chen +4
Instruction-following is a critical capability of Large Language Models (LLMs). While existing works primarily focus on assessing how well LLMs adhere to user instructions, they of…