2 papers
cs.CL2025
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
Xiang Li, Wenyue Hua, Kaijie Zhu +8
Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal understanding, yet their reasoning abilities remain underexplored. Existing benchmarks…
cs.CL2025
SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
Zekun Li, Shinda Huang, Jiangtian Wang +8
As language agents increasingly automate critical tasks, their ability to follow domain-specific standard operating procedures (SOPs), policies, and constraints when taking actions…