4 papers
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li, Yimin Liu, Wenbo Chen +75
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to m…
System-Level Analysis of Module Uncertainty Quantification in the Autonomy Pipeline
Sampada Deglurkar, Haotian Shen, Anish Muthali +5
Modern autonomous systems with machine learning components often use uncertainty quantification to help produce assurances about system operation. However, there is a lack of conse…
Generative Sequential Notification Optimization via Multi-Objective Decision Transformers
Borja Ocejo, Ruofan Wang, Ke Liu +7
Notifications are an important communication channel for delivering timely and relevant information. Optimizing their delivery involves addressing complex sequential decision-makin…
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents
LeCheng Zhang, Yuanshi Wang, Haotian Shen +1
The Da Vinci Code, a game of logical deduction and imperfect information, presents unique challenges for artificial intelligence, demanding nuanced reasoning beyond simple pattern…