6 papers
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel +185
Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…
Adversarial Hubness Detector: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems
Idan Habler, Vineeth Sai Narajala, Stav Koren +2
Retrieval-Augmented Generation (RAG) systems are essential to contemporary AI applications, allowing large language models to obtain external knowledge via vector similarity search…
CoPE: A Small Language Model for Steerable and Scalable Content Labeling
Samidh Chakrabarti, David Willner, Kevin Klyman +3
This paper details the methodology behind CoPE, a policy-steerable small language model capable of fast and accurate content labeling. We present a novel training curricula called…
Cisco Integrated AI Security and Safety Framework Report
Amy Chang, Tiffany Saade, Sanket Mendapara +2
Artificial intelligence (AI) systems are being readily and rapidly adopted, increasingly permeating critical domains: from consumer platforms and enterprise software to networked s…
A2AS: Agentic AI Runtime Security and Self-Defense
Eugene Neelou, Ivan Novikov, Max Moroz +15
The A2AS framework is introduced as a security layer for AI agents and LLM-powered applications, similar to how HTTPS secures HTTP. A2AS enforces certified behavior, activates mode…
User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
Jennifer King, Kevin Klyman, Emily Capstick +2
Hundreds of millions of people now regularly interact with large language models via chatbots. Model developers are eager to acquire new sources of high-quality training data as th…