Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity
Samuel Jacob Chacko, James Hugglestone, Chashi Mahiul Islam +1
Agent Skills, structured packages of procedural knowledge loaded into an LLM agent at inference time, are widely reported to improve task pass rates by an average of 16.2~percentag…
cs.AI2026
Numerical Instability and Chaos: Quantifying the Unpredictability of Large Language Models
Chashi Mahiul Islam, Alan Villarreal, Mao Nishino +2
As Large Language Models (LLMs) are increasingly integrated into agentic workflows, their unpredictability stemming from numerical instability has emerged as a critical reliability…