5 papers
SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills
Xinze Chen, Chi Zhang, Ping Ji +1
Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored. We pr…
What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files
Chi Zhang, Yimin Liu, Xinze Chen +1
Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) agents to reuse procedures beyo…
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li, Yimin Liu, Wenbo Chen +75
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to m…
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Xiangyi Li, Kyoung Whan Choe, Yimin Liu +12
Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evaluating them on live services is r…
DeepDR: an integrated deep-learning model web server for drug repositioning
Shuting Jin, Yi Jiang, Yimin Liu +5
Background: Identifying new indications for approved drugs is a complex and time-consuming process that requires extensive knowledge of pharmacology, clinical data, and advanced co…