3 papers
cs.AI2025
AI Benchmark Democratization and Carpentry
Gregor von Laszewski, Wesley Brewer, Jeyan Thiyagalingam +28
Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring d…
cs.CR2025
Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world
Tobin South, Subramanya Nagabhushanaradhya, Ayesha Dissanayaka +18
The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand fo…
cs.SE2025
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
Sean McGregor, Victor Lu, Vassil Tashev +8
Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by…