Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
Ruida Wang, Jerry Huang, Pengcheng Wang +3
Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recent advances in LLMs' agentic…
cs.AI2026
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li, Yimin Liu, Wenbo Chen +75
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to m…