Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
Pei-An Chen, Yong-Ching Liang, Jia-Fong Yeh +4
Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. However, existing methods usually…
cs.AI2025
OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models
Yuhe Liu, Changhua Pei, Longlong Xu +13
Information Technology (IT) Operations (Ops), particularly Artificial Intelligence for IT Operations (AIOps), is the guarantee for maintaining the orderly and stable operation of e…