5 papers
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
Aris Hofmann, Inge Vejsbjerg, Dhaval Salwala +1
We present Auto-BenchmarkCard, a workflow for generating validated descriptions of AI benchmarks. Benchmark documentation is often incomplete or inconsistent, making it difficult t…
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly +17
The rapid evolution of generative AI has expanded the breadth of risks associated with AI systems. While various taxonomies and frameworks exist to classify these risks, the lack o…
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
Anna Sokol, Elizabeth Daly, Michael Hind +4
Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation method…
Usage Governance Advisor: From Intent to AI Governance
Elizabeth M. Daly, Sean Rooney, Seshu Tirupathi +9
Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deploy…
Granite Guardian
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia +20
We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with…