7 papers
AI Steerability 360: A Toolkit for Steering Large Language Models
Erik Miehling, Karthikeyan Natesan Ramamurthy, Praveen Venkateswaran +10
The AI Steerability 360 toolkit is an extensible, open-source Python library for steering LLMs. Steering abstractions are designed around four model control surfaces: input (modifi…
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
Aris Hofmann, Inge Vejsbjerg, Dhaval Salwala +1
We present Auto-BenchmarkCard, a workflow for generating validated descriptions of AI benchmarks. Benchmark documentation is often incomplete or inconsistent, making it difficult t…
ICX360: In-Context eXplainability 360 Toolkit
Dennis Wei, Ronny Luss, Xiaomeng Hu +6
Large Language Models (LLMs) have become ubiquitous in everyday life and are entering higher-stakes applications ranging from summarizing meeting transcripts to answering doctors'…
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly +17
The rapid evolution of generative AI has expanded the breadth of risks associated with AI systems. While various taxonomies and frameworks exist to classify these risks, the lack o…
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models
Seshu Tirupathi, Dhaval Salwala, Elizabeth Daly +1
As Large Language Models (LLMs) continue to be increasingly applied across various domains, their widespread adoption necessitates rigorous monitoring to prevent unintended negativ…
Humble AI in the real-world: the case of algorithmic hiring
Rahul Nair, Inge Vejsbjerg, Elizabeth Daly +2
Humble AI (Knowles et al., 2023) argues for cautiousness in AI development and deployments through scepticism (accounting for limitations of statistical learning), curiosity (accou…