Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
Kazem Faghih, Yize Cheng, Shoumik Saha +3
Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in differe…
cs.AI2026
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
Shoumik Saha, Kazem Faghih, Soheil Feizi
Autonomous AI agents increasingly extend their capabilities through Agent Skills: modular filesystem packages whose SKILL.md files describe when and how agents should use them. Whi…