2 papers
cs.CR2026
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
Yiran Wu, Mauricio Velazco, Andrew Zhao +9
We present ExCyTIn-Bench, the first benchmark to Evaluate an LLM agent X on the task of Cyber Threat Investigation through security questions derived from investigation graphs. Rea…
cs.AI2025
Dynamic Context-Aware Prompt Recommendation for Domain-Specific AI Applications
Xinye Tang, Haijun Zhai, Chaitanya Belwal +3
LLM-powered applications are highly susceptible to the quality of user prompts, and crafting high-quality prompts can often be challenging especially for domain-specific applicatio…