4 papers
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
Aris Hofmann, Inge Vejsbjerg, Dhaval Salwala +1
We present Auto-BenchmarkCard, a workflow for generating validated descriptions of AI benchmarks. Benchmark documentation is often incomplete or inconsistent, making it difficult t…
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly +17
The rapid evolution of generative AI has expanded the breadth of risks associated with AI systems. While various taxonomies and frameworks exist to classify these risks, the lack o…
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models
Seshu Tirupathi, Dhaval Salwala, Elizabeth Daly +1
As Large Language Models (LLMs) continue to be increasingly applied across various domains, their widespread adoption necessitates rigorous monitoring to prevent unintended negativ…
Usage Governance Advisor: From Intent to AI Governance
Elizabeth M. Daly, Sean Rooney, Seshu Tirupathi +9
Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deploy…