From the 1 of 6 linked papers with an AI index.
6 papers
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Paul Kassianik, Blaine Nelson, Yaron Singer
The paper proposes a cost‑aware evaluation framework for language‑model security agents, measuring both offensive CTF performance and defensive SOC investigation efficiency by fixi…
FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines
Paul Kassianik, Baturay Saglam, Huaibo Zhao +4
Multi-step LLM pipelines fail through interactions among retrieval, reasoning, and formatting steps, so prompt-only optimization can miss bottlenecks in the chain. We present Fully…
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
Zhuoran Yang, Ed Li, Jianliang He +18
We present Foundation-Sec-8B-Reasoning, the first open-source native reasoning model for cybersecurity. Built upon our previously released Foundation-Sec-8B base model (derived fro…
Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
Baturay Saglam, Paul Kassianik, Blaine Nelson +3
Understanding the latent space geometry of large language models (LLMs) is key to interpreting their behavior and improving alignment. Yet it remains unclear to what extent LLMs li…
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
Sajana Weerawardhena, Paul Kassianik, Blaine Nelson +14
Large language models (LLMs) have shown remarkable success across many domains, yet their integration into cybersecurity applications remains limited due to a lack of general-purpo…
Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
Paul Kassianik, Baturay Saglam, Alexander Chen +15
As transformer-based large language models (LLMs) increasingly permeate society, they have revolutionized domains such as software engineering, creative writing, and digital arts.…