4 papers
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
Aris Hofmann, Inge Vejsbjerg, Dhaval Salwala +1
We present Auto-BenchmarkCard, a workflow for generating validated descriptions of AI benchmarks. Benchmark documentation is often incomplete or inconsistent, making it difficult t…
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models
Seshu Tirupathi, Dhaval Salwala, Elizabeth Daly +1
As Large Language Models (LLMs) continue to be increasingly applied across various domains, their widespread adoption necessitates rigorous monitoring to prevent unintended negativ…
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly +17
The rapid evolution of generative AI has expanded the breadth of risks associated with AI systems. While various taxonomies and frameworks exist to classify these risks, the lack o…
Usage Governance Advisor: From Intent to AI Governance
Elizabeth M. Daly, Sean Rooney, Seshu Tirupathi +9
Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deploy…